Skip to content

Commit f983042

Browse files
aroma+studio: food-only corpus (42 heads, open-gov cited) + formulation export
Data / food-safety pass (this is a flavor app — non-food odorants are removed from the training corpus, not merely flagged): - Remove isovanillin, habanolide (no food clearance), methyleugenol (delisted from 21 CFR 172.515 in 2018), and estragole/methyl chavicol (prohibited as an added flavouring in the EU, Reg. 1334/2008 Annex III). - Reclaim musky, spicy, vanilla on food-authorised molecules, and add new anise and meaty heads -> 42 aroma heads (fresh dropped, 0.639 borderline). - food_safe_supplement.csv: 38 molecules, each cited to an OPEN-GOVERNMENT register only — EU/GB Union List (data.food.gov.uk, OGL v3) or US 21 CFR / FDA SAF (public domain). Commercial compilations (Good Scents / Leffingwell / FEMA library) are finding-aids for the public FL number, never a source. Accurate food-use labelling (was imprecise): - The reference is FDA Substances-Added-to-Food (broader than GRAS) + EU flavourings, so status now reads 'listed as an authorised food flavouring (EU FL 07.142)' / '... FDA Substances-Added-to-Food reference' with the specific authority, never a blanket 'GRAS'. UI badges: GRAS -> Food-listed. Formulation export (parity with the flavor-card export): - /api/recipe_card PNG endpoint; Export CSV + Export card buttons on a designed recipe; design_recipe_csv MCP tool (15 tools) and skill 'design --csv'; e2e test covering both downloads. Fixes: - design_recipe / gap-analysis food-safe filter matched _gras_status TEXT, which broke when the status string changed -> now checks InChIKey skeleton membership in the reference set directly (robust). - food_safe_supplement read with dtype=str so FL numbers keep leading zeros. Sources fully cited in docs/SOURCES.md, docs/DATA-SOURCES.md, and the web UI sources section (EU Union List / OGL added; finding-aid vs source distinction). Signed-off-by: Austin L. <86896075+rvnminers-A-and-N@users.noreply.github.com>
1 parent 0725f70 commit f983042

18 files changed

Lines changed: 467 additions & 63 deletions

.gitignore

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -64,3 +64,4 @@ Thumbs.db
6464
.vscode/
6565
*.swp
6666
!training/aroma_supplement.csv
67+
!training/food_safe_supplement.csv

README.md

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -17,15 +17,15 @@ running entirely on hardware you own.
1717
<p align="center">
1818
<img src="docs/assets/flavor-map.png" alt="Flavor-space map in 3D on MW × logP × TPSA axes, colored by taste and aroma" width="900">
1919
</p>
20-
<p align="center"><sub>The interactive flavor-space map in 3D on real <b>MW × logP × TPSA</b> axes, colored by <b>taste &amp; aroma</b> — every one of the 41 aroma + 6 taste classes labelled.</sub></p>
20+
<p align="center"><sub>The interactive flavor-space map in 3D on real <b>MW × logP × TPSA</b> axes, colored by <b>taste &amp; aroma</b> — every one of the 42 aroma + 6 taste classes labelled.</sub></p>
2121

2222
> **8,393 unique molecules** across the open datasets · **taste + aroma** prediction from
2323
> structure · a **flavor library** (start from a flavor → its character-impact molecule) and
2424
> **flavor designer** (pick your notes → best food-safe molecules + drop-in swaps) · an
2525
> interactive 2D/3D **flavor-space map** · **2D & 3D** structure views. All on
2626
> commercial-clean public data.
2727
>
28-
> Per set (unique molecules): taste training **3,845** · aroma training **1046** · odor
28+
> Per set (unique molecules): taste training **3,845** · aroma training **1060** · odor
2929
> corpus **2,255** · documented taste **676** · GRAS reference **2,781** · sweetness
3030
> intensity **316** · curated character-impact flavors **95 flavors / 77 molecules** · public-domain aroma supplement **102 associations**.
3131
> Every one of the 8,393 is enriched with names + measured properties from public-domain PubChem.
@@ -59,7 +59,7 @@ is tagged by how it was derived**, so nothing reads as more certain than its sou
5959
**tasteless** (RandomForests on fingerprint + physicochemical features), plus a
6060
sweetness-**intensity** regressor. Sour and salty *also* keep a transparent chemistry
6161
rule (acid group / alkali-salt) as a deterministic cross-check alongside the model.
62-
- **Aroma****41 odor-descriptor heads** (citrus, floral, minty, almond, fatty,
62+
- **Aroma****42 odor-descriptor heads** (citrus, floral, minty, almond, fatty,
6363
petroleum, earthy, medicinal, sulfurous, camphor, fruity, fishy, garlic, ethereal,
6464
ammoniacal, pungent, pine, rose, rancid, alcoholic, woody, green, grassy, putrid)
6565
trained on **public-domain** HSDB odor text + curated character-impact facts, surfaced
@@ -154,7 +154,7 @@ Flavormancer ships as two editions of one method:
154154
| Edition | Commercial | Academic / open-source *(coming soon)* |
155155
| License | Apache-2.0 | open-source, **research / NonCommercial** |
156156
| Data | commercial-clean open data only | adds research odor datasets with **NonCommercial** terms |
157-
| Aroma | **41 presence/absence descriptor heads ship** (public-domain HSDB); scored **intensity** is trained on your data or a licensed set (PMP 2001) | full open model incl. **intensity** (research odor data) |
157+
| Aroma | **42 presence/absence descriptor heads ship** (public-domain HSDB); scored **intensity** is trained on your data or a licensed set (PMP 2001) | full open model incl. **intensity** (research odor data) |
158158
| Use | free to use, sell, run on-prem | research, teaching, advancing the method |
159159

160160
The split is deliberate. The richest aroma data is licensed for research only, so

docs/AROMA-AUDIT.md

Lines changed: 24 additions & 17 deletions
Original file line numberDiff line numberDiff line change
@@ -1,24 +1,31 @@
11
# Aroma-head audit — where more data buys more heads
22

3-
> **Update — supplement applied.** `build_aroma_supplement.py` added a curated **public-domain**
4-
> character-impact set (`aroma_supplement.csv`, resolved via PubChem) for the sparse descriptors
5-
> below. Result: **24 → 41 heads****17 new**: `coconut` (0.89), `nutty` (0.94), `caramel` (0.91),
6-
> `winey` (0.90), `onion` (0.87), `honey` (0.80), `herbal` (0.80), `vanilla` (0.93), `buttery` (0.87),
7-
> `balsamic` (0.93), `smoky` (0.99), `cinnamon` (1.00), `spicy` (0.71), `banana` (0.95),
8-
> `fresh` (0.73), `musky` (1.00), `vegetable` (0.85) — and `grassy` kept above the
9-
> bar (0.80) with its classic green-leaf volatiles. `coconut` is now *predictable* (γ-nonalactone →
10-
> coconut 1.0), which closes its earlier data-gate. `spicy`/`fresh` are borderline (0.71–0.73).
3+
> **Update — supplement applied, then made food-only.** `build_aroma_supplement.py` adds a curated
4+
> **public-domain** character-impact set (`aroma_supplement.csv`, resolved via PubChem) for the sparse
5+
> descriptors below. Result: **24 → 42 heads.** New heads include `coconut`, `nutty`, `caramel`,
6+
> `winey`, `onion`, `honey`, `herbal`, `vanilla`, `buttery`, `balsamic`, `smoky`, `cinnamon` (1.00),
7+
> `spicy` (0.78), `banana`, `musky` (1.00), `vegetable`, and — added in the food-safety pass —
8+
> `anise` (0.80) and `meaty` (0.89). `coconut` is now *predictable* (γ-nonalactone → coconut 1.0).
119
> **Only `sweet`-odor (0.637) remains unlearnable** even with 185 examples — a genuine representation
12-
> limit (revisit with a GNN).
10+
> limit (revisit with a GNN). `fresh` was evaluated but **dropped** (0.639, high-variance/borderline).
1311
>
14-
> **Small-n caveat (honest):** the *highest*-AUROC new heads are small, structurally-homogeneous
15-
> classes — e.g. `cinnamon` and `musky` reach CV-AUROC **~1.00** only because their ~10 positives
16-
> are one scaffold (cinnamaldehyde family; macrocyclic musks) the fingerprint separates trivially.
17-
> That is memorising a scaffold,
18-
> not a superb general model: these heads are **narrow and high-variance** (like `grassy`/`coffee`)
19-
> and will sharpen — or get honestly re-scored — with more diverse examples. Treat AUROC on n≈10
20-
> heads as indicative, not gospel. The numbers
21-
> below describe the **pre-supplement** baseline + the standing plan.
12+
> **Food-only + open-government provenance (this is a flavor app).** Non-food odorants were **removed
13+
> from the corpus**, not merely flagged: `isovanillin` and `habanolide` (no food clearance),
14+
> `methyleugenol` (delisted from 21 CFR 172.515 in 2018), and `estragole`/`methyl chavicol` (prohibited
15+
> as an added flavouring in the EU, Reg. 1334/2008 Annex III). Heads that leaned on them (`musky`,
16+
> `vanilla`, `spicy`) were **re-based on food-authorised molecules** and retrained above the bar. Every
17+
> food-clearance is cited to an **open-government register** — the EU/GB Union List (`data.food.gov.uk`,
18+
> OGL v3) or US 21 CFR / FDA SAF (public domain) — in `food_safe_supplement.csv`; **no commercial
19+
> compilation** (Good Scents / Leffingwell / FEMA library) is a source, only a finding-aid for the
20+
> public FL number.
21+
>
22+
> **Small-n caveat (honest):** the *highest*-AUROC heads are small, structurally-homogeneous classes —
23+
> e.g. `cinnamon` and `musky` reach CV-AUROC **~1.00** only because their ~10–12 positives are one
24+
> scaffold family (cinnamaldehyde esters; macrocyclic musk lactones/ketones) the fingerprint separates
25+
> trivially. That is memorising a scaffold, not a superb general model: these heads are **narrow and
26+
> high-variance** and will sharpen — or get honestly re-scored — with more diverse examples. Treat
27+
> AUROC on n≈10 heads as indicative, not gospel. The numbers below describe the **pre-supplement**
28+
> baseline + the standing plan.
2229
2330

2431
An honest look at the aroma model: which odor descriptors we can predict from structure today,

docs/AROMA.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -18,7 +18,7 @@ enumerable via the annotations API. Pipeline:
1818
physicochemical descriptor block** (`chemfeatures.py`: MW, logP, TPSA, H-bond donors/
1919
acceptors, rotatable bonds, ring counts, fraction sp3, heteroatoms — the global properties
2020
bare bits miss). After broadening the vocabulary, folding the curated `flavors.csv` molecules
21-
in (825 → **1034** molecules, incl. a curated public-domain supplement), and adding those features, **41 heads clear CV-AUROC ≥ 0.70**
21+
in (825 → **1060** molecules, incl. a curated public-domain supplement), and adding those features, **42 heads clear CV-AUROC ≥ 0.70**
2222
(each score is that descriptor's own held-out AUROC): medicinal 0.97, ammoniacal 0.96,
2323
petroleum 0.96, citrus 0.94, camphor 0.93, almond 0.92, fatty 0.89, minty 0.88, ethereal 0.87,
2424
fruity 0.87, fishy 0.86, sulfurous 0.85, garlic 0.84, floral 0.83, pungent 0.81, earthy 0.73.

docs/CAPABILITIES.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -32,7 +32,7 @@ by how it's derived, and nothing claims more certainty than its source supports.
3232
- Multitaste — fires when 2+ taste heads are high (**trained**-derived).
3333
- Known-taste ground truth — verified labels override predictions (**lookup**).
3434

35-
**Aroma****41 odor-descriptor heads ship** (**trained**)
35+
**Aroma****42 odor-descriptor heads ship** (**trained**)
3636
- Presence/absence per descriptor (citrus, floral, minty, almond, fatty, petroleum, earthy,
3737
medicinal, sulfurous, camphor, fruity, fishy, garlic, ethereal, ammoniacal, pungent) —
3838
RandomForests on fingerprint + physicochemical features, over public-domain HSDB odor text +

docs/DATA-SOURCES.md

Lines changed: 6 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -12,7 +12,9 @@ What feeds the models, what each source unlocks, its license, and how to get it.
1212
> CC-BY-4.0** (confirmed on the Zenodo dataset deposit) and **SweetenersDB MIT**
1313
> (confirmed on `github.com/chemosim-lab/SweetenersDB`) for taste + sweetness; **PubChem
1414
> / HSDB / CAMEO / Haz-Map / Tox21 / FDA-SAF** are US-government public domain for
15-
> odor text, properties, tox assays, and GRAS. **No NonCommercial data is in the
15+
> odor text, properties, tox assays, and GRAS; the **EU/GB flavourings Union List**
16+
> (`data.food.gov.uk`, **Open Government Licence v3** — commercial reuse permitted) supplies
17+
> food-clearance facts for a handful of molecules absent from FDA-SAF. **No NonCommercial data is in the
1618
> commercial edition** — FlavorDB (CC-BY-NC-SA), the Leffingwell/GoodScents **GS-LF**
1719
> aroma set (NonCommercial), the FartDB composite, and PlantMolecularTasteDB are
1820
> **excluded** and reserved for the **academic edition** (or a licensed source such as
@@ -31,7 +33,9 @@ What feeds the models, what each source unlocks, its license, and how to get it.
3133
| **SweetenersDB v2.0** | the sweetness-**intensity** regressor | **MIT** | ✅ in use | direct CSV from `github.com/chemosim-lab/SweetenersDB` (`SweetenersDB_v2.0.csv`, 316 cmpds, `logSw` column). R²≈0.82. |
3234
| **Pyrfume — Leffingwell/GoodScents (GS-LF)** | the **aroma model** (OpenPOM) — #17/#18 | **RESTRICTED** |**excluded** | Use restrictions (John Leffingwell & Google; GoodScents/Arctander/Flavornet © Datu Inc.). **GS-LF = OpenPOM's GoodScents + Leffingwell *combined* set (~4,983 cmpds — rich precisely because it merges both), NonCommercial.** Excluded from the commercial edition entirely; it's the **academic edition's** aroma data. *Distinct from* **PMP 2001** — Leffingwell's *paid, commercially-licensable* product — which the commercial edition could use if licensed. |
3335
| **Pyrfume — open academic sets** | aroma model, **clean but small** | CC-BY / open-access (verify each) | candidates | `keller_2016` (BMC, CC-BY, ~480 cmpds), `snitz_2013` (PLOS, CC-BY), etc. Confirm each data deposit's license; combine + harmonize descriptor vocabularies for volume. |
34-
| **FEMA GRAS / FDA SAF** | GRAS cross-reference + dosing/OAV lookups | gov public / FEMA | ⬜ to get | FDA "Substances Added to Food" (public domain) → `gras_reference.parquet`; FEMA use-level PDFs → `properties.parquet`. **Verify every scraped dosing number.** |
36+
| **FDA SAF (Substances Added to Food)** | GRAS/food-ingredient cross-reference (the defensive "recognized food ingredient?" signal) | US-gov **public domain** | ✅ in use | FDA "Substances Added to Food" inventory (`build_gras_reference.py``gras_reference.parquet`, CAS→InChIKey via PubChem). |
37+
| **EU/GB flavourings Union List** | food-clearance for a few aroma-supplement molecules absent from the SAF crawl (`food_safe_supplement.csv`) | **Open Government Licence v3 / EU law (Reg. 1334/2008 Annex I)** — commercial reuse permitted | ✅ in use | Queried at `data.food.gov.uk/regulated-products/flavouring_authorisations/<FL>`; each entry's chemical name + status confirmed. **8 molecules** added (EU FL numbers / 21 CFR 172.515), unioned into the same GRAS signal. Regulatory facts are non-copyrightable (Feist). **NOT sourced from Good Scents / FEMA flavor-library / Leffingwell** — those are finding aids only, never the citation. |
38+
| **Excluded — non-food odorants** |||**removed from corpus** | `isovanillin` (no EU FL / CFR clearance; "not for flavor use") and `habanolide` (fragrance-only musk, no corroborated food clearance) were **dropped from the aroma supplement** — this is a flavor app, so molecules that aren't food ingredients are kept out of the training data, not merely flagged. Dropping them retired the `musky` head (fragrance-leaning) and `vanilla` was re-based on food-safe molecules. |
3539

3640
## Commercial data-completion plan (the data-gated features)
3741

docs/HOW-IT-WORKS.md

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -62,7 +62,7 @@ features (trained on SweetenersDB).
6262

6363
---
6464

65-
## 3. Aroma — 41 descriptor heads from public odor text
65+
## 3. Aroma — 42 descriptor heads from public odor text
6666

6767
Odor is the hard, licensed part of this field — most rich odor datasets are NonCommercial. Our
6868
clean route:
@@ -75,7 +75,7 @@ clean route:
7575
`minty`), producing a presence/absence label per descriptor. We also fold in the curated
7676
character-impact facts from `flavors.csv`.
7777
3. **`train_aroma.py`** trains one RandomForest per descriptor (on fingerprint + descriptors) and
78-
**keeps only the heads that clear CV-AUROC ≥ 0.70** — an honest bar. **41 heads** survive
78+
**keeps only the heads that clear CV-AUROC ≥ 0.70** — an honest bar. **42 heads** survive
7979
(citrus, floral, minty, almond, fatty, petroleum, earthy, medicinal, sulfurous, camphor,
8080
fruity, fishy, garlic, ethereal, ammoniacal, pungent, pine, rose, rancid, alcoholic,
8181
woody, green, grassy, putrid), 0.71–0.98. The lower-population heads (≈10–20 documented

docs/SOURCES.md

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -56,8 +56,19 @@ record of provenance, not legal advice. Get an IP/OSS-license review before ship
5656
**Safety / regulatory / properties**
5757
- **FDA "Substances Added to Food" (SAF)** — the GRAS / recognized-food-ingredient
5858
cross-check. **US-government work, public domain.** In use via `build_gras_reference.py`.
59+
- **EU / GB flavourings Union List** — Regulation (EC) 1334/2008 Annex I, published as the
60+
authorised-flavourings register at **`data.food.gov.uk/regulated-products/flavouring_authorisations`**.
61+
Used to confirm food-clearance (by **FL number**) for ~26 aroma-supplement molecules absent from
62+
the FDA SAF crawl (`food_safe_supplement.csv`). **Open Government Licence v3 — commercial reuse
63+
permitted.** Regulatory facts (a molecule's authorised status + FL number) are non-copyrightable
64+
(*Feist v. Rural*, 1991). **Finding-aid note:** commercial compilations (The Good Scents Company,
65+
the FEMA flavor library, Leffingwell) were used *only* to locate a candidate FL number, which was
66+
then **confirmed on the open-government register**; no content from those compilations is copied,
67+
redistributed, or shipped. **12** further molecules are cleared via **US 21 CFR 172.515 / FDA SAF**
68+
(public domain), flagged US-only where they carry no EU/GB entry.
5969
- **FEMA usual/maximum use levels** — for *quantitative* dosing. Published in FEMA's
6070
copyrighted GRAS papers — **not freely available in bulk**; data-gated (customer/licensed).
71+
We cite FEMA *numbers* only as public regulatory identifiers, never FEMA's compiled dosing text.
6172
- **EU declarable fragrance/flavor allergen annex** — the labeling flags. (EU regulation.)
6273
- **PubChem** (NIH/NCBI) — used three ways: name/CAS↔SMILES↔CID resolution; experimental
6374
**boiling point / vapor pressure** (`build_properties.py`); and CAS→InChIKey for GRAS

mcp-server/README.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -14,7 +14,7 @@ MCP client ──stdio──▶ flavormancer-mcp ──HTTP──▶ Flavormance
1414

1515
## Tools
1616

17-
Full coverage of Flavormancer's capabilities — 13 tools.
17+
Full coverage of Flavormancer's capabilities — 15 tools.
1818

1919
**Single molecule**
2020

mcp-server/server.py

Lines changed: 22 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -296,6 +296,28 @@ def design_recipe(flavors: list[str] = [], notes: list[str] = [], food_safe: boo
296296
return _design_recipe(flavors, notes, food_safe)
297297

298298

299+
@mcp.tool()
300+
def design_recipe_csv(flavors: list[str] = [], notes: list[str] = [], food_safe: bool = True) -> str:
301+
"""Same as design_recipe, but returns the recipe as a CSV bench sheet (a string) instead of
302+
JSON — parity with the workbench's 'Export CSV' and the skill CLI's `design --csv`. Columns:
303+
ingredient, smiles, ppm, volatility, carries, dose_basis. Hand this straight to a formulator
304+
or drop it into a spreadsheet."""
305+
import csv as _csv
306+
import io as _io
307+
out = _design_recipe(flavors, notes, food_safe)
308+
if isinstance(out, dict) and out.get("error"):
309+
return "error," + str(out["error"])
310+
rec = (out or {}).get("recipe") or []
311+
buf = _io.StringIO()
312+
buf.write("# Flavormancer formulation - directional starting recipe (tune on the bench)\n")
313+
w = _csv.writer(buf)
314+
w.writerow(["ingredient", "smiles", "ppm", "volatility", "carries", "dose_basis"])
315+
for i in rec:
316+
w.writerow([i.get("name", ""), i.get("smiles", ""), i.get("ppm", ""),
317+
i.get("volatility", ""), "; ".join(i.get("carries") or []), i.get("dose_basis", "")])
318+
return buf.getvalue()
319+
320+
299321
@mcp.tool()
300322
def screen_mixture(ingredients: list[str], processes: list[str] = []) -> dict:
301323
"""Screen a mixture of ingredients for DOCUMENTED food hazards (e.g. benzoate + ascorbate

0 commit comments

Comments
 (0)