Skip to content

Pull the whole public-domain odor corpus + fix name display - #81

Merged
rvnminers-A-and-N merged 1 commit into
mainfrom
odor-corpus-and-names
Jul 4, 2026
Merged

rvnminers-A-and-N merged 1 commit into
mainfrom
odor-corpus-and-names

Conversation

@rvnminers-A-and-N

Copy link
Copy Markdown
Collaborator

--pubchem-all: pulls the entire public-domain PubChem odor set via the annotations API (inline text+source+CID, all pages), filtered to HSDB/Haz-Map/CAMEO, batched CID→InChIKey. odor_notes.parquet ~190 → ~2,260 molecules — the training corpus for the descriptor model.

Name fix: flattened-structure Titles are often systematic (cinnamaldehyde → '3-Phenylprop-2-Enal'), which looked like the IUPAC name doubled. Now the typed name is used as the common name, and the display never repeats the same string.

build_odor_notes.py gains --pubchem-all: pulls the ENTIRE public-domain PubChem odor
set via the annotations API (odor text + source + CID inline, all pages), filtered to
HSDB/Haz-Map/CAMEO, resolving CIDs to InChIKeys in batches. ~20 requests instead of a
3k-molecule crawl; odor_notes.parquet goes from ~190 to ~2,260 molecules.

Name fix: PubChem's Title for a flattened structure is often the systematic name
(cinnamaldehyde -> '3-Phenylprop-2-Enal'), which read as the IUPAC name shown twice.
Now api/names uses the user's typed name as the common name when they searched by
name, and the workbench never renders common when it equals the IUPAC string.

Signed-off-by: Austin L. <86896075+rvnminers-A-and-N@users.noreply.github.com>
@rvnminers-A-and-N
rvnminers-A-and-N merged commit 3cccd56 into main Jul 4, 2026
3 checks passed
@rvnminers-A-and-N
rvnminers-A-and-N deleted the odor-corpus-and-names branch July 4, 2026 14:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant