v2.1 shipped: YOLO26n produce localization plus a DINOv3-S/16 24-class freshness classifier with open-world rejection. Local Streamlit app, honest cluster-disjoint metrics, Open Images negative evaluation, and an external KTH GroceryStoreDataset type benchmark.
Demo · Quickstart · Eval report · Training notebooks · PRD
End-to-end run through the Specimen Lab page: image upload → YOLO26n produce localization → scene-box filtering → DINOv3-S/16 freshness labels where evidence is strong → confidence + source badges. Unsupported uploads abstain instead of being forced into the nearest produce class.
The GIF is hosted in the profile repo to keep this repo lean. To run it yourself, follow Quickstart below.
| Metric | Value |
|---|---|
| Classifier macro F1 (24-class) | 0.9478 |
| Classifier top-1 accuracy | 0.9490 |
| KTH external type accuracy | 0.8995 |
| KTH external sample count | 955 |
| Detector mAP@50 | 0.8693 |
| Detector mAP@50–95 | 0.8190 |
| Open Images negative false accept | 0.0873 |
| Open Images positive retention | 0.5951 |
Macro F1 is the headline because top-1 accuracy hides minority-class failure under the 41 : 1 class imbalance baked into the source dataset. The KTH row is type-only external evidence; KTH has no fresh/rotten labels, so it is not mixed into the canonical 24-class freshness metric. The Open Images rows measure open-world detector behavior on existing validated web-style data. Numbers come from eval_report.json, produced by notebooks/kaggle_05_evaluate_v2.ipynb.
The flow below is the same pipeline as the figure above, rendered as a flowchart so the routing logic is precise and reviewable.
flowchart LR
UP["📷 Uploaded image"]
YOLO["YOLO26n · produce-only<br/>boxes · confidence"]
CROP["DINOv3-S/16<br/>crop classifier · 24-class"]
FULL["DINOv3-S/16<br/>full-image sanity check"]
POST["Open-world filters<br/>drop scene boxes · reject weak full-frame boxes"]
DET["✅ Detections<br/>(label · freshness · box · conf)"]
PARTIAL["Apple / n_a<br/>type known · freshness uncertain"]
UNK["🔶 unknown / n_a"]
UP --> YOLO
YOLO -->|"candidate boxes"| POST
POST -->|"item-level boxes"| CROP
POST -->|"no usable produce evidence"| UNK
UP --> FULL
CROP -->|"crop/full-image agree or full abstains"| DET
CROP -->|"same type · freshness split"| PARTIAL
CROP -->|"type disagreement or weak evidence"| UNK
YOLO -->|"no boxes"| UNK
The 24-class label space is {12 produce types} × {fresh, rotten}. In v2 the detector has no type or freshness opinion; DINOv3 is the single authority for labels, applied to each crop, with the full image used as an out-of-distribution sanity check. Runtime filters now reject no-box uploads and weak near-full-frame detections instead of asking the closed-set classifier to guess.
git clone https://github.com/Abdulrahman-Elsmmany/freshguard-vision.git
cd freshguard-vision
uv sync
uv run python scripts/download_artifacts.py
uv run streamlit run app.pyRequires Python 3.12, uv, and local PyTorch model weights downloaded from a GitHub Release. Inference is local PyTorch only — no cloud APIs, no ONNX, no remote services.
The checkpoints are release assets, not git files. The app expects:
| File | Purpose | Size |
|---|---|---|
yolo26n_produce_v2_1.pt |
one-class produce detector with Open Images negative hardening | ~5 MB |
dinov3_vits16_food_freshness_v2.pt |
DINOv3-S/16 24-class classifier | ~83 MB |
The download script fetches both into artifacts/:
uv run python scripts/download_artifacts.pyManual download is also fine: open the latest GitHub Release, download both
.pt files, and place them in artifacts/. The app will stay in a clear
"models not ready" state until both files are present.
Best-performing classes (classifier macro F1):
| Class | F1 | Support |
|---|---|---|
bitter_gourd_fresh |
1.000 | 48 |
bitter_gourd_rotten |
1.000 | 54 |
strawberry_fresh |
1.000 | 147 |
banana_rotten |
0.995 | 555 |
strawberry_rotten |
0.989 | 90 |
Weakest:
| Class | F1 | Support |
|---|---|---|
carrot_rotten |
0.855 | 381 |
orange_rotten |
0.865 | 452 |
bellpepper_fresh |
0.866 | 353 |
bellpepper_rotten |
0.886 | 356 |
tomato_rotten |
0.897 | 477 |
The weakest v2 classes are still above 0.85 F1, a large lift over the v1 okra_* floor. Most remaining confusion sits near fine-grained visual boundaries and the fresh ↔ rotten boundary within a single produce type. Full breakdown: eval_report.md.
Every step runs on Kaggle. The six v2 notebooks under notebooks/ chain their outputs through Kaggle Datasets — each notebook publishes a dataset that the next one attaches as input.
| # | Notebook | Kaggle inputs | Accelerator | Save output as |
|---|---|---|---|---|
| 0 | kaggle_00_fetch_official_sources_v2.ipynb |
none; Internet on | none | freshguard-official-sources-v2-1 |
| 1 | kaggle_01_dataset_audit_v2.ipynb |
ulnnproject/food-freshness-dataset, freshguard-official-sources-v2 |
none | freshguard-v2-splits |
| 2 | kaggle_02_prepare_detector_data_v2.ipynb |
ulnnproject/food-freshness-dataset, freshguard-v2-splits, freshguard-official-sources-v2-1 |
none | freshguard-v2-1-detector-data |
| 3 | kaggle_03_train_detector_v2.ipynb |
freshguard-v2-1-detector-data |
T4 x2 | freshguard-v2-1-detector-artifacts |
| 4 | kaggle_04_train_classifier_dinov3_v2.ipynb |
freshguard-v2-splits, ulnnproject/food-freshness-dataset, freshguard-official-sources-v2 |
T4 x2 | freshguard-v2-classifier-artifacts |
| 5 | kaggle_05_evaluate_v2.ipynb |
splits, detector data, detector artifact, classifier artifact, official sources, Food Freshness, optional five-apple dataset | P100 or T4 | freshguard-v2-1-eval |
The v1 notebooks remain in the directory as the shipped v0.2.0 evidence trail.
freshguard-vision/
├── app.py · Streamlit entrypoint
├── configs/inference.toml · Runtime config (thresholds, paths)
├── .streamlit/config.toml · Theme tokens
├── notebooks/ · Kaggle training/evaluation notebooks
├── scripts/
│ └── download_artifacts.py · Pulls .pt weights from GitHub Release
├── src/freshness/
│ ├── inference/ · Pipeline + detector + classifier wrappers
│ ├── ui/ · Streamlit pages, theming, components
│ ├── utils/ · Image I/O, label normalization
│ ├── config.py
│ └── constants.py · 24-class label space, Latin binomials
├── PRD.md · Product brief — scope, goals, non-goals
├── eval_report.md · Headline metrics, human-readable
└── eval_report.json · Same metrics, machine-readable
The .pt checkpoints are large and don't belong in git. v1 outputs are archived locally under deprecated/ and remain out of version control. The v2 runtime expects:
yolo26n_produce_v2_1.pt— produce-only detector with Open Images negative hardeningdinov3_vits16_food_freshness_v2.pt— DINOv3-S/16 classifier
scripts/download_artifacts.py resolves the latest release via the GitHub API and drops both files into artifacts/ automatically.
- The weakest v2 classifier classes are still imperfect:
carrot_rotten(F1 0.855),orange_rotten(0.865), andbellpepper_fresh(0.866). - KTH is type-only external evidence. It has official grocery image splits but no fresh/rotten labels, so it cannot validate the 24-class freshness contract by itself.
- Detector supervision is mixed: Food Freshness and KTH contribute full-image bootstrap boxes, while Open Images contributes official object boxes plus v2.1 negative/background examples with empty YOLO labels. Runtime scene-box suppression reduces the visible impact, but more instance-level boxes remain the next detector-quality lever.
- Open Images positive retention is the next detector target: v2.1 measures a low negative false-accept rate (0.0873), but retains 0.5951 of Open Images positive images at the current runtime threshold.
- Out-of-distribution images (cluttered scenes, novel varieties, non-produce subjects) trigger the detector/open-world gates, classifier abstain stack, crop/full-image disagreement guard, or
unknown— the system would rather refuse than confidently hallucinate. - Freshness can be uncertain even when type is clear. Same-produce fresh/rotten ties are surfaced as
<produce> / n_ainstead of pretending the binary freshness decision is settled. - Binary freshness only. No shelf-life forecast, no "medium" / partial-ripeness label.
Python 3.12 · uv · PyTorch 2.8 · Ultralytics 8.3 / YOLO26n · timm / DINOv3 · Streamlit 1.56 · Grounding DINO (training only)
Local Streamlit demo · v2.1 release v0.3.1 · single-author project · MIT licensed.
Abdulrahman Elsmmany · eng.elsmmany@gmail.com · linkedin
MIT License · © 2026

