pino-source-code / TODO.md
Matthew Ford
docs: TODO ledger β€” authoritative GPU readout, CPU demoted to superseded, A/B task marked complete
8460eb6
|
Raw
History Blame Contribute Delete
12.5 kB
# PINO research TODO
This file is the execution ledger for the representation-validation work. A checked
box means that the frozen artifact, machine-readable report, and tests (where
applicable) are committed. Results are recorded even when they are negative.
## Representation validation sequence
- [x] Run frozen-feature ablations comparing Morgan, real POM, and POM + physics
on descriptor, odor-threshold, substantivity, and substitution tasks.
- Real POM beat Morgan on every predictive task (5-fold paired bootstrap):
substantivity rho 0.679 vs 0.392 (delta +0.287, 95%CI [+0.206,+0.368]);
odor_threshold rho 0.601 vs 0.528 (delta +0.073, [+0.033,+0.111]);
descriptor macroF1 0.291 vs 0.137 (delta +0.154, [+0.095,+0.214]).
- Fixed a registry_props() bug (missing odor_threshold_ug_m3) that had silently
emptied the threshold task. Results: `artifacts/representation_ablation/`.
- [x] Create and freeze a discordant substitution-triplet benchmark.
- 8 genuine triplets admitted: independent supplier preference provenance +
verified structural discordance (Tanimoto distractor > preferred).
- EXPANDED 8 -> 20 (2026-07-17): `load_trade_name_structures` now also mines
`material_profiles` (molequles) for trade-name -> SMILES, resolving
proprietary trade materials (Isobutavan, Canthoxal, Neofolione, Indocolore,
Centifolether, Romandolide, ...). All 12 new triplets carry authoritative
Fraterworks provenance. Coverage still bounded by structural identity;
expand with expert-rated perceptual preferences for publication scale.
`data/benchmarks/substitution_triplets/`.
- [x] Freeze a prospective formula benchmark containing 30--50 deliberately
designed formulas (COMPLETE; see design decisions below).
- DECISION: evaluated in-silico (model scoring/retrieval), no physical
compounding. OPTION HELD OPEN: a community physical-compounding arm may be
appended later; formulas are therefore restricted to common hobbyist-available
aroma chemicals so the same set works for both.
- Palette: 51 synthetics resolved to SMILES via material_profiles (molequles),
32 already in the real-POM asset (extend POM to the rest).
- DONE: 40 preregistered formulas frozen; intended family profiles sequestered
(`labels.sequestered.json`, access=evaluation_only).
- [x] Train the next PIMT candidate only after all three benchmark artifacts above
are frozen and their manifests pass the training-readiness gate.
- Gate `ready_for_training=true` (all 5 checks). Two-arm representation A/B
trained (6 epochs, CPU): Morgan fallback vs genuine 256-dim OpenPOM via
`--structural-source`. Checkpoints + model card pushed to
`mattbitzesty/pino-pimt-representation-ab` (public, verified HTTP 200).
- [x] Evaluate the candidate once on the frozen benchmarks and report all outcomes,
including negative and inconclusive results.
- AUTHORITATIVE (GPU, 20 epochs, n=20 triplets): val_total morgan 0.3959 vs
openpom 0.3600; triplets morgan 0.30 vs openpom 0.45; prospective family
cosine morgan 0.5378 vs openpom 0.6121. OpenPOM ahead on all three; both
arms below chance on triplets (reported negative). `artifacts/frozen_eval_*_gpu.json`.
- CPU readout (6 epochs, n=8; superseded): val_total 0.4797/0.4722; triplets
0.25/0.50; prospective 0.5257/0.5349. Directional only. Evaluator
`scripts/evaluate_frozen_benchmarks.py`.
- [x] Re-run the two-arm A/B longer on HF GPU (20 epochs, T4) against the expanded
20-triplet benchmark; upload checkpoints + `frozen_eval_*_gpu.json` to
`mattbitzesty/pino-pimt-representation-ab`. COMPLETE (job 6a59cdae): OpenPOM
direction confirmed at 20 epochs; consolidated stats `paper_stats_bundle.json`.
## Molequles -> training link (2026-07-17)
The training descriptor matrix (generate_targets_v3.build_master_matrix) never read
the molequles-enriched material_profiles; ingredients absent from the registry fell
through every tier to a zero 138-D vector.
- [x] Added Tier 3.5: material_profiles odor text -> existing Tier-2 tokenizer.
Ingredient-instance coverage on the repaired corpus 90.7% -> 98.5%
(uncovered instances 3353 -> 531). Commit d6969e8.
- [ ] NEW DIRECTIVE (open): train on the FULL molequles tag vocabulary (~600 tags
at >=10 occurrences) rather than only the curated 138-Pyrfume categories, which
may be more insightful. DECISION PENDING on target-space architecture:
138-dim single-label coarse basis vs ~600-dim multi-label tag basis. This is a
paper-defining architectural fork; see README "Open design decisions".
- [x] Close the residual 531-instance gap (15 CAS with no odor text anywhere) via
curated perfumer-tier labels β€” coverage now 99.80% (36,049/36,122 instances;
73-instance name-only naturals floor).
- [x] Test suite green: fixed pre-existing failures (draft_engine None-MW crash;
stale gate-failure test rewritten to hermetic tmp-dir + live-READY regression
guard). 112 passed, 4 skipped.
## Completion evidence
- [x] Frozen-feature ablation manifest and report (COMPLETE; real POM wins):
`artifacts/representation_ablation/`
- [x] Discordant triplet manifest and report (8 admitted; honest coverage note):
`data/benchmarks/substitution_triplets/`
- [x] Prospective formula protocol and blinded evaluation package (40 frozen,
labels sequestered): `data/benchmarks/prospective_formulas/`
- [x] Training-readiness gate report (`READY_FOR_TRAINING`; all 5 checks green):
`artifacts/training_readiness.json`
- [x] Paper data exporter (`artifacts/paper_data.yaml`, 49 macros) + figures
(`figures/*.pdf`): `scripts/export_paper_data.py`, `scripts/generate_paper_figures.py`
- [x] v11.6 substitution-corpus repair (Molequles enrichment; gate now READY):
`artifacts/pimt_v11_6_corrected_coverage_baseline.json`
## v11.6 substitution-corpus repair (2026-07-17)
Data-side lever, executed per the v11.5 pre-training investigation's recommended
sequence. Rights-cleared Molequles corpus (3,090 molecule groups) used as an
independent enrichment source β€” never for eval-label construction.
- Profile table enriched: `material_profiles_v11_6.jsonl` (5,738 profiles).
+730 descriptor distributions, +1,094 substantivity external priors,
regulatory_status 0β†’1,750, IFRA cat-4 limits 0β†’223. Fill-null-only policy;
measured/modelled PINO fields never overwritten.
- Identity adjudication: 7 self-resolved target/answer rows excluded
(v11_sub_0014 surfaced only after enrichment); Neofolione(111-79-5) /
Habanolide(111879-80-2) merge repaired.
- Answer recovery: 9 Molequles-backed profiles admitted + 6 commercial-alias
bridges (Cyclopidene, Neobergamate Forte, Wood 49, Orris absolute, Vetiver,
Rum Ether), all with independent-identity evidence.
- Evidence floor added to scorer: rows with <2 mutually-observed components
flagged `low_evidence`; headline reported both all-evaluable and
trustworthy-only.
- Corrected baseline: rankable one-to-one 16β†’29 (+81%), median answer rank
1022.5β†’253, best rank 8, top-5 hits still 0. Retrieval features (not ranking
weights) remain the binding constraint. (Note: this baseline predates the
training-readiness gate, which has since gone READY on the benchmark checks.)
All three former prerequisites are now met (real POM asset, 20-triplet benchmark,
40 prospective formulas), so the gate reads READY.
## Compliant-repair pipeline (2026-07-17)
Salvaged 58 banned-material formulas (Fraterworks) by replacing restricted
ingredients with compliant alternatives, and in doing so produced a second,
provenance-clean substitution benchmark.
- Scanned 131 Fraterworks formulas; 58 contained banned/IFRA-restricted
materials (musk ketone, oakmoss, lilial, lyral, musk xylene).
- Hard compliance filter on the v11.6 substitution engine: no banned-for-banned
swaps, no same-class swaps, IFRA-unrestricted preferred. Self-check passed,
zero violations, all 76 repairs `ifra_unrestricted`.
- Two repair kinds separated: 60 true substitutions + 16 identity corrections
(source rows mislabeled compliant "Veramoss" under banned oakmoss CAS
90028-68-5; corrected to Evernyl 4707-47-5).
- Artifacts: `data/repaired_formulas_v11_6.jsonl` (58 original+repaired pairs)
and `data/benchmarks/repair_pairs/` (76-pair substitution benchmark +
manifest).
- Surfaced a real profile-table defect: 241 cross-identity alias collisions
(e.g. "Veramoss" aliased under Oakmoss absolute). Repair pipeline routes
around it via CAS-aware resolution; a central alias-layer cleanup is now a
named follow-up.
This is the second independent substitution corpus (after the 83-row supplier
set), and the first to be generated by the project's own engine with a
compliance guarantee.
## Full-corpus compliant repair (2026-07-17, supersedes Fraterworks-only run)
Extended the repair to every formula corpus with per-source adapters, and a
repair vs scan-only policy split.
- Repair corpora (literature/historical, rewritten): fraterworks, appell,
poucher, literature_flat β†’ **106 repair pairs**, 91 formulas repaired (87
fully). 90 true substitutions + 16 identity corrections. `appell_flat`
excluded as a 100% duplicate of `appell`.
- Compliance: zero violations, all 106 repairs `ifra_unrestricted`.
- Substitutes: only 6 distinct compliant materials used (Ambrettolide 77Γ—,
Veramoss/Evernyl 16Γ—, Lilyflore 11Γ—, Florhydral 10Γ—, Fauxmoss 7Γ—).
- **Scan-only finding (not rewritten):** the live training corpus carries
**309 banned-material occurrences** β€” tgsc_demo 209 (167/492 formulas),
empirical v9+wisemoor 100 (64/5,708). By class: lilial 141, lyral 75,
musk ketone 52, oakmoss 28, musk xylene 13. This is a training-data-quality
exposure: the model is being trained on formulas that are not legally
manufacturable. Repairing these is a training-data decision (it rewrites
ground-truth compositions) and needs an explicit call before running.
- Artifacts: `data/repaired_formulas_v11_6_all.jsonl`,
`data/benchmarks/repair_pairs/repair_pairs_all.jsonl` + `manifest_all.json`,
report `artifacts/pimt_v11_6_compliant_repair_all.json`. Tests 111/111.
## Trainable-corpus repair + applicability research (2026-07-17)
Aggressive option: repair the live training corpus in place, with a hard
guarantee that no trained row is a formula/target mismatch or carries a banned
material. Executed as a chain of provenance-tagged passes.
- **Compliance repair** of tgsc_demo (167 formulas) + empirical (212 formulas):
460 pairs, all `ifra_unrestricted`, zero violations.
- **Target regeneration** for empirical repaired rows (targets are computed
from composition, so a repair without re-derivation is a mismatch). 160
regenerated cleanly; 52 could not be simulated β†’ rolled back to original
rather than ship a mismatch.
- **Identity repair** (typo'd CAS): Evernyl 4707-47-3β†’4707-47-5, Sandalore
65113-99-5β†’65113-99-7, Geranyl acetate 164-09-4β†’105-87-3, Ethyl Linalool
39255-35-5β†’10339-55-6, Vetyveryl acetate 117-98-6β†’62563-80-8. Applied only
when the name's canonical CAS is unambiguous.
- **Applicability research + unblock:** the residual blockers are structures
ugropy's Dortmund-UNIFAC ILP cannot assign (gamma-pyrones: maltol/ethyl
maltol, diphenyl oxide, helional) plus 14 natural oils absent from the
19-entry expansion table. Verified the VLE already tolerates empty groups
(ideal Ξ³=1.0) without wrecking the physics β†’ unblocked those rows with a
LABELED ideal-solution fallback (never mixed with real UNIFAC output).
Researched and added approximate GC constituent profiles for 9 blocked
naturals (peppermint, spearmint, basil, roman camomile, armoise, bay laurel,
tarragon, taget, copaiba) and expanded them into single-CAS constituents.
- **Compliance leak caught and closed:** an intermediate pass simulated 23 rows
without re-applying the banned substitution, leaving banned CAS in trainable
rows. Fixed; final scan: **0 banned CAS in any trainable row**.
**Final empirical state: 5,678 / 5,708 trainable (99.47%).** 177 banned-
material rows recovered into the trainable set. 30 rows remain blocked β€” all
natural oils (costus, basil variants, tarragon, camomile, copaiba, etc.) whose
exact GC compositions are not in any local source; blocked rows are isolated
(repair_status set), not trained. Remaining unblock = sourcing/authoring
constituent profiles for those 14 naturals. Tests 111/111 green.