# PINO research TODO This file is the execution ledger for the representation-validation work. A checked box means that the frozen artifact, machine-readable report, and tests (where applicable) are committed. Results are recorded even when they are negative. ## Representation validation sequence - [x] Run frozen-feature ablations comparing Morgan, real POM, and POM + physics on descriptor, odor-threshold, substantivity, and substitution tasks. - Real POM beat Morgan on every predictive task (5-fold paired bootstrap): substantivity rho 0.679 vs 0.392 (delta +0.287, 95%CI [+0.206,+0.368]); odor_threshold rho 0.601 vs 0.528 (delta +0.073, [+0.033,+0.111]); descriptor macroF1 0.291 vs 0.137 (delta +0.154, [+0.095,+0.214]). - Fixed a registry_props() bug (missing odor_threshold_ug_m3) that had silently emptied the threshold task. Results: `artifacts/representation_ablation/`. - [x] Create and freeze a discordant substitution-triplet benchmark. - 8 genuine triplets admitted: independent supplier preference provenance + verified structural discordance (Tanimoto distractor > preferred). - EXPANDED 8 -> 20 (2026-07-17): `load_trade_name_structures` now also mines `material_profiles` (molequles) for trade-name -> SMILES, resolving proprietary trade materials (Isobutavan, Canthoxal, Neofolione, Indocolore, Centifolether, Romandolide, ...). All 12 new triplets carry authoritative Fraterworks provenance. Coverage still bounded by structural identity; expand with expert-rated perceptual preferences for publication scale. `data/benchmarks/substitution_triplets/`. - [x] Freeze a prospective formula benchmark containing 30--50 deliberately designed formulas (COMPLETE; see design decisions below). - DECISION: evaluated in-silico (model scoring/retrieval), no physical compounding. OPTION HELD OPEN: a community physical-compounding arm may be appended later; formulas are therefore restricted to common hobbyist-available aroma chemicals so the same set works for both. - Palette: 51 synthetics resolved to SMILES via material_profiles (molequles), 32 already in the real-POM asset (extend POM to the rest). - DONE: 40 preregistered formulas frozen; intended family profiles sequestered (`labels.sequestered.json`, access=evaluation_only). - [x] Train the next PIMT candidate only after all three benchmark artifacts above are frozen and their manifests pass the training-readiness gate. - Gate `ready_for_training=true` (all 5 checks). Two-arm representation A/B trained (6 epochs, CPU): Morgan fallback vs genuine 256-dim OpenPOM via `--structural-source`. Checkpoints + model card pushed to `mattbitzesty/pino-pimt-representation-ab` (public, verified HTTP 200). - [x] Evaluate the candidate once on the frozen benchmarks and report all outcomes, including negative and inconclusive results. - AUTHORITATIVE (GPU, 20 epochs, n=20 triplets): val_total morgan 0.3959 vs openpom 0.3600; triplets morgan 0.30 vs openpom 0.45; prospective family cosine morgan 0.5378 vs openpom 0.6121. OpenPOM ahead on all three; both arms below chance on triplets (reported negative). `artifacts/frozen_eval_*_gpu.json`. - CPU readout (6 epochs, n=8; superseded): val_total 0.4797/0.4722; triplets 0.25/0.50; prospective 0.5257/0.5349. Directional only. Evaluator `scripts/evaluate_frozen_benchmarks.py`. - [x] Re-run the two-arm A/B longer on HF GPU (20 epochs, T4) against the expanded 20-triplet benchmark; upload checkpoints + `frozen_eval_*_gpu.json` to `mattbitzesty/pino-pimt-representation-ab`. COMPLETE (job 6a59cdae): OpenPOM direction confirmed at 20 epochs; consolidated stats `paper_stats_bundle.json`. ## Molequles -> training link (2026-07-17) The training descriptor matrix (generate_targets_v3.build_master_matrix) never read the molequles-enriched material_profiles; ingredients absent from the registry fell through every tier to a zero 138-D vector. - [x] Added Tier 3.5: material_profiles odor text -> existing Tier-2 tokenizer. Ingredient-instance coverage on the repaired corpus 90.7% -> 98.5% (uncovered instances 3353 -> 531). Commit d6969e8. - [ ] NEW DIRECTIVE (open): train on the FULL molequles tag vocabulary (~600 tags at >=10 occurrences) rather than only the curated 138-Pyrfume categories, which may be more insightful. DECISION PENDING on target-space architecture: 138-dim single-label coarse basis vs ~600-dim multi-label tag basis. This is a paper-defining architectural fork; see README "Open design decisions". - [x] Close the residual 531-instance gap (15 CAS with no odor text anywhere) via curated perfumer-tier labels — coverage now 99.80% (36,049/36,122 instances; 73-instance name-only naturals floor). - [x] Test suite green: fixed pre-existing failures (draft_engine None-MW crash; stale gate-failure test rewritten to hermetic tmp-dir + live-READY regression guard). 112 passed, 4 skipped. ## Completion evidence - [x] Frozen-feature ablation manifest and report (COMPLETE; real POM wins): `artifacts/representation_ablation/` - [x] Discordant triplet manifest and report (8 admitted; honest coverage note): `data/benchmarks/substitution_triplets/` - [x] Prospective formula protocol and blinded evaluation package (40 frozen, labels sequestered): `data/benchmarks/prospective_formulas/` - [x] Training-readiness gate report (`READY_FOR_TRAINING`; all 5 checks green): `artifacts/training_readiness.json` - [x] Paper data exporter (`artifacts/paper_data.yaml`, 49 macros) + figures (`figures/*.pdf`): `scripts/export_paper_data.py`, `scripts/generate_paper_figures.py` - [x] v11.6 substitution-corpus repair (Molequles enrichment; gate now READY): `artifacts/pimt_v11_6_corrected_coverage_baseline.json` ## v11.6 substitution-corpus repair (2026-07-17) Data-side lever, executed per the v11.5 pre-training investigation's recommended sequence. Rights-cleared Molequles corpus (3,090 molecule groups) used as an independent enrichment source — never for eval-label construction. - Profile table enriched: `material_profiles_v11_6.jsonl` (5,738 profiles). +730 descriptor distributions, +1,094 substantivity external priors, regulatory_status 0→1,750, IFRA cat-4 limits 0→223. Fill-null-only policy; measured/modelled PINO fields never overwritten. - Identity adjudication: 7 self-resolved target/answer rows excluded (v11_sub_0014 surfaced only after enrichment); Neofolione(111-79-5) / Habanolide(111879-80-2) merge repaired. - Answer recovery: 9 Molequles-backed profiles admitted + 6 commercial-alias bridges (Cyclopidene, Neobergamate Forte, Wood 49, Orris absolute, Vetiver, Rum Ether), all with independent-identity evidence. - Evidence floor added to scorer: rows with <2 mutually-observed components flagged `low_evidence`; headline reported both all-evaluable and trustworthy-only. - Corrected baseline: rankable one-to-one 16→29 (+81%), median answer rank 1022.5→253, best rank 8, top-5 hits still 0. Retrieval features (not ranking weights) remain the binding constraint. (Note: this baseline predates the training-readiness gate, which has since gone READY on the benchmark checks.) All three former prerequisites are now met (real POM asset, 20-triplet benchmark, 40 prospective formulas), so the gate reads READY. ## Compliant-repair pipeline (2026-07-17) Salvaged 58 banned-material formulas (Fraterworks) by replacing restricted ingredients with compliant alternatives, and in doing so produced a second, provenance-clean substitution benchmark. - Scanned 131 Fraterworks formulas; 58 contained banned/IFRA-restricted materials (musk ketone, oakmoss, lilial, lyral, musk xylene). - Hard compliance filter on the v11.6 substitution engine: no banned-for-banned swaps, no same-class swaps, IFRA-unrestricted preferred. Self-check passed, zero violations, all 76 repairs `ifra_unrestricted`. - Two repair kinds separated: 60 true substitutions + 16 identity corrections (source rows mislabeled compliant "Veramoss" under banned oakmoss CAS 90028-68-5; corrected to Evernyl 4707-47-5). - Artifacts: `data/repaired_formulas_v11_6.jsonl` (58 original+repaired pairs) and `data/benchmarks/repair_pairs/` (76-pair substitution benchmark + manifest). - Surfaced a real profile-table defect: 241 cross-identity alias collisions (e.g. "Veramoss" aliased under Oakmoss absolute). Repair pipeline routes around it via CAS-aware resolution; a central alias-layer cleanup is now a named follow-up. This is the second independent substitution corpus (after the 83-row supplier set), and the first to be generated by the project's own engine with a compliance guarantee. ## Full-corpus compliant repair (2026-07-17, supersedes Fraterworks-only run) Extended the repair to every formula corpus with per-source adapters, and a repair vs scan-only policy split. - Repair corpora (literature/historical, rewritten): fraterworks, appell, poucher, literature_flat → **106 repair pairs**, 91 formulas repaired (87 fully). 90 true substitutions + 16 identity corrections. `appell_flat` excluded as a 100% duplicate of `appell`. - Compliance: zero violations, all 106 repairs `ifra_unrestricted`. - Substitutes: only 6 distinct compliant materials used (Ambrettolide 77×, Veramoss/Evernyl 16×, Lilyflore 11×, Florhydral 10×, Fauxmoss 7×). - **Scan-only finding (not rewritten):** the live training corpus carries **309 banned-material occurrences** — tgsc_demo 209 (167/492 formulas), empirical v9+wisemoor 100 (64/5,708). By class: lilial 141, lyral 75, musk ketone 52, oakmoss 28, musk xylene 13. This is a training-data-quality exposure: the model is being trained on formulas that are not legally manufacturable. Repairing these is a training-data decision (it rewrites ground-truth compositions) and needs an explicit call before running. - Artifacts: `data/repaired_formulas_v11_6_all.jsonl`, `data/benchmarks/repair_pairs/repair_pairs_all.jsonl` + `manifest_all.json`, report `artifacts/pimt_v11_6_compliant_repair_all.json`. Tests 111/111. ## Trainable-corpus repair + applicability research (2026-07-17) Aggressive option: repair the live training corpus in place, with a hard guarantee that no trained row is a formula/target mismatch or carries a banned material. Executed as a chain of provenance-tagged passes. - **Compliance repair** of tgsc_demo (167 formulas) + empirical (212 formulas): 460 pairs, all `ifra_unrestricted`, zero violations. - **Target regeneration** for empirical repaired rows (targets are computed from composition, so a repair without re-derivation is a mismatch). 160 regenerated cleanly; 52 could not be simulated → rolled back to original rather than ship a mismatch. - **Identity repair** (typo'd CAS): Evernyl 4707-47-3→4707-47-5, Sandalore 65113-99-5→65113-99-7, Geranyl acetate 164-09-4→105-87-3, Ethyl Linalool 39255-35-5→10339-55-6, Vetyveryl acetate 117-98-6→62563-80-8. Applied only when the name's canonical CAS is unambiguous. - **Applicability research + unblock:** the residual blockers are structures ugropy's Dortmund-UNIFAC ILP cannot assign (gamma-pyrones: maltol/ethyl maltol, diphenyl oxide, helional) plus 14 natural oils absent from the 19-entry expansion table. Verified the VLE already tolerates empty groups (ideal γ=1.0) without wrecking the physics → unblocked those rows with a LABELED ideal-solution fallback (never mixed with real UNIFAC output). Researched and added approximate GC constituent profiles for 9 blocked naturals (peppermint, spearmint, basil, roman camomile, armoise, bay laurel, tarragon, taget, copaiba) and expanded them into single-CAS constituents. - **Compliance leak caught and closed:** an intermediate pass simulated 23 rows without re-applying the banned substitution, leaving banned CAS in trainable rows. Fixed; final scan: **0 banned CAS in any trainable row**. **Final empirical state: 5,678 / 5,708 trainable (99.47%).** 177 banned- material rows recovered into the trainable set. 30 rows remain blocked — all natural oils (costus, basil variants, tarragon, camomile, copaiba, etc.) whose exact GC compositions are not in any local source; blocked rows are isolated (repair_status set), not trained. Remaining unblock = sourcing/authoring constituent profiles for those 14 naturals. Tests 111/111 green.