pino-source-code / TODO.md
Matthew Ford
docs: TODO ledger β€” authoritative GPU readout, CPU demoted to superseded, A/B task marked complete
8460eb6
|
Raw
History Blame Contribute Delete
12.5 kB

PINO research TODO

This file is the execution ledger for the representation-validation work. A checked box means that the frozen artifact, machine-readable report, and tests (where applicable) are committed. Results are recorded even when they are negative.

Representation validation sequence

  • Run frozen-feature ablations comparing Morgan, real POM, and POM + physics on descriptor, odor-threshold, substantivity, and substitution tasks.
    • Real POM beat Morgan on every predictive task (5-fold paired bootstrap): substantivity rho 0.679 vs 0.392 (delta +0.287, 95%CI [+0.206,+0.368]); odor_threshold rho 0.601 vs 0.528 (delta +0.073, [+0.033,+0.111]); descriptor macroF1 0.291 vs 0.137 (delta +0.154, [+0.095,+0.214]).
    • Fixed a registry_props() bug (missing odor_threshold_ug_m3) that had silently emptied the threshold task. Results: artifacts/representation_ablation/.
  • Create and freeze a discordant substitution-triplet benchmark.
    • 8 genuine triplets admitted: independent supplier preference provenance + verified structural discordance (Tanimoto distractor > preferred).
    • EXPANDED 8 -> 20 (2026-07-17): load_trade_name_structures now also mines material_profiles (molequles) for trade-name -> SMILES, resolving proprietary trade materials (Isobutavan, Canthoxal, Neofolione, Indocolore, Centifolether, Romandolide, ...). All 12 new triplets carry authoritative Fraterworks provenance. Coverage still bounded by structural identity; expand with expert-rated perceptual preferences for publication scale. data/benchmarks/substitution_triplets/.
  • Freeze a prospective formula benchmark containing 30--50 deliberately designed formulas (COMPLETE; see design decisions below).
    • DECISION: evaluated in-silico (model scoring/retrieval), no physical compounding. OPTION HELD OPEN: a community physical-compounding arm may be appended later; formulas are therefore restricted to common hobbyist-available aroma chemicals so the same set works for both.
    • Palette: 51 synthetics resolved to SMILES via material_profiles (molequles), 32 already in the real-POM asset (extend POM to the rest).
    • DONE: 40 preregistered formulas frozen; intended family profiles sequestered (labels.sequestered.json, access=evaluation_only).
  • Train the next PIMT candidate only after all three benchmark artifacts above are frozen and their manifests pass the training-readiness gate.
    • Gate ready_for_training=true (all 5 checks). Two-arm representation A/B trained (6 epochs, CPU): Morgan fallback vs genuine 256-dim OpenPOM via --structural-source. Checkpoints + model card pushed to mattbitzesty/pino-pimt-representation-ab (public, verified HTTP 200).
  • Evaluate the candidate once on the frozen benchmarks and report all outcomes, including negative and inconclusive results.
    • AUTHORITATIVE (GPU, 20 epochs, n=20 triplets): val_total morgan 0.3959 vs openpom 0.3600; triplets morgan 0.30 vs openpom 0.45; prospective family cosine morgan 0.5378 vs openpom 0.6121. OpenPOM ahead on all three; both arms below chance on triplets (reported negative). artifacts/frozen_eval_*_gpu.json.
    • CPU readout (6 epochs, n=8; superseded): val_total 0.4797/0.4722; triplets 0.25/0.50; prospective 0.5257/0.5349. Directional only. Evaluator scripts/evaluate_frozen_benchmarks.py.
  • Re-run the two-arm A/B longer on HF GPU (20 epochs, T4) against the expanded 20-triplet benchmark; upload checkpoints + frozen_eval_*_gpu.json to mattbitzesty/pino-pimt-representation-ab. COMPLETE (job 6a59cdae): OpenPOM direction confirmed at 20 epochs; consolidated stats paper_stats_bundle.json.

Molequles -> training link (2026-07-17)

The training descriptor matrix (generate_targets_v3.build_master_matrix) never read the molequles-enriched material_profiles; ingredients absent from the registry fell through every tier to a zero 138-D vector.

  • Added Tier 3.5: material_profiles odor text -> existing Tier-2 tokenizer. Ingredient-instance coverage on the repaired corpus 90.7% -> 98.5% (uncovered instances 3353 -> 531). Commit d6969e8.
  • NEW DIRECTIVE (open): train on the FULL molequles tag vocabulary (~600 tags at >=10 occurrences) rather than only the curated 138-Pyrfume categories, which may be more insightful. DECISION PENDING on target-space architecture: 138-dim single-label coarse basis vs ~600-dim multi-label tag basis. This is a paper-defining architectural fork; see README "Open design decisions".
  • Close the residual 531-instance gap (15 CAS with no odor text anywhere) via curated perfumer-tier labels β€” coverage now 99.80% (36,049/36,122 instances; 73-instance name-only naturals floor).
  • Test suite green: fixed pre-existing failures (draft_engine None-MW crash; stale gate-failure test rewritten to hermetic tmp-dir + live-READY regression guard). 112 passed, 4 skipped.

Completion evidence

  • Frozen-feature ablation manifest and report (COMPLETE; real POM wins): artifacts/representation_ablation/
  • Discordant triplet manifest and report (8 admitted; honest coverage note): data/benchmarks/substitution_triplets/
  • Prospective formula protocol and blinded evaluation package (40 frozen, labels sequestered): data/benchmarks/prospective_formulas/
  • Training-readiness gate report (READY_FOR_TRAINING; all 5 checks green): artifacts/training_readiness.json
  • Paper data exporter (artifacts/paper_data.yaml, 49 macros) + figures (figures/*.pdf): scripts/export_paper_data.py, scripts/generate_paper_figures.py
  • v11.6 substitution-corpus repair (Molequles enrichment; gate now READY): artifacts/pimt_v11_6_corrected_coverage_baseline.json

v11.6 substitution-corpus repair (2026-07-17)

Data-side lever, executed per the v11.5 pre-training investigation's recommended sequence. Rights-cleared Molequles corpus (3,090 molecule groups) used as an independent enrichment source β€” never for eval-label construction.

  • Profile table enriched: material_profiles_v11_6.jsonl (5,738 profiles). +730 descriptor distributions, +1,094 substantivity external priors, regulatory_status 0β†’1,750, IFRA cat-4 limits 0β†’223. Fill-null-only policy; measured/modelled PINO fields never overwritten.
  • Identity adjudication: 7 self-resolved target/answer rows excluded (v11_sub_0014 surfaced only after enrichment); Neofolione(111-79-5) / Habanolide(111879-80-2) merge repaired.
  • Answer recovery: 9 Molequles-backed profiles admitted + 6 commercial-alias bridges (Cyclopidene, Neobergamate Forte, Wood 49, Orris absolute, Vetiver, Rum Ether), all with independent-identity evidence.
  • Evidence floor added to scorer: rows with <2 mutually-observed components flagged low_evidence; headline reported both all-evaluable and trustworthy-only.
  • Corrected baseline: rankable one-to-one 16β†’29 (+81%), median answer rank 1022.5β†’253, best rank 8, top-5 hits still 0. Retrieval features (not ranking weights) remain the binding constraint. (Note: this baseline predates the training-readiness gate, which has since gone READY on the benchmark checks.)

All three former prerequisites are now met (real POM asset, 20-triplet benchmark, 40 prospective formulas), so the gate reads READY.

Compliant-repair pipeline (2026-07-17)

Salvaged 58 banned-material formulas (Fraterworks) by replacing restricted ingredients with compliant alternatives, and in doing so produced a second, provenance-clean substitution benchmark.

  • Scanned 131 Fraterworks formulas; 58 contained banned/IFRA-restricted materials (musk ketone, oakmoss, lilial, lyral, musk xylene).
  • Hard compliance filter on the v11.6 substitution engine: no banned-for-banned swaps, no same-class swaps, IFRA-unrestricted preferred. Self-check passed, zero violations, all 76 repairs ifra_unrestricted.
  • Two repair kinds separated: 60 true substitutions + 16 identity corrections (source rows mislabeled compliant "Veramoss" under banned oakmoss CAS 90028-68-5; corrected to Evernyl 4707-47-5).
  • Artifacts: data/repaired_formulas_v11_6.jsonl (58 original+repaired pairs) and data/benchmarks/repair_pairs/ (76-pair substitution benchmark + manifest).
  • Surfaced a real profile-table defect: 241 cross-identity alias collisions (e.g. "Veramoss" aliased under Oakmoss absolute). Repair pipeline routes around it via CAS-aware resolution; a central alias-layer cleanup is now a named follow-up.

This is the second independent substitution corpus (after the 83-row supplier set), and the first to be generated by the project's own engine with a compliance guarantee.

Full-corpus compliant repair (2026-07-17, supersedes Fraterworks-only run)

Extended the repair to every formula corpus with per-source adapters, and a repair vs scan-only policy split.

  • Repair corpora (literature/historical, rewritten): fraterworks, appell, poucher, literature_flat β†’ 106 repair pairs, 91 formulas repaired (87 fully). 90 true substitutions + 16 identity corrections. appell_flat excluded as a 100% duplicate of appell.
  • Compliance: zero violations, all 106 repairs ifra_unrestricted.
  • Substitutes: only 6 distinct compliant materials used (Ambrettolide 77Γ—, Veramoss/Evernyl 16Γ—, Lilyflore 11Γ—, Florhydral 10Γ—, Fauxmoss 7Γ—).
  • Scan-only finding (not rewritten): the live training corpus carries 309 banned-material occurrences β€” tgsc_demo 209 (167/492 formulas), empirical v9+wisemoor 100 (64/5,708). By class: lilial 141, lyral 75, musk ketone 52, oakmoss 28, musk xylene 13. This is a training-data-quality exposure: the model is being trained on formulas that are not legally manufacturable. Repairing these is a training-data decision (it rewrites ground-truth compositions) and needs an explicit call before running.
  • Artifacts: data/repaired_formulas_v11_6_all.jsonl, data/benchmarks/repair_pairs/repair_pairs_all.jsonl + manifest_all.json, report artifacts/pimt_v11_6_compliant_repair_all.json. Tests 111/111.

Trainable-corpus repair + applicability research (2026-07-17)

Aggressive option: repair the live training corpus in place, with a hard guarantee that no trained row is a formula/target mismatch or carries a banned material. Executed as a chain of provenance-tagged passes.

  • Compliance repair of tgsc_demo (167 formulas) + empirical (212 formulas): 460 pairs, all ifra_unrestricted, zero violations.
  • Target regeneration for empirical repaired rows (targets are computed from composition, so a repair without re-derivation is a mismatch). 160 regenerated cleanly; 52 could not be simulated β†’ rolled back to original rather than ship a mismatch.
  • Identity repair (typo'd CAS): Evernyl 4707-47-3β†’4707-47-5, Sandalore 65113-99-5β†’65113-99-7, Geranyl acetate 164-09-4β†’105-87-3, Ethyl Linalool 39255-35-5β†’10339-55-6, Vetyveryl acetate 117-98-6β†’62563-80-8. Applied only when the name's canonical CAS is unambiguous.
  • Applicability research + unblock: the residual blockers are structures ugropy's Dortmund-UNIFAC ILP cannot assign (gamma-pyrones: maltol/ethyl maltol, diphenyl oxide, helional) plus 14 natural oils absent from the 19-entry expansion table. Verified the VLE already tolerates empty groups (ideal Ξ³=1.0) without wrecking the physics β†’ unblocked those rows with a LABELED ideal-solution fallback (never mixed with real UNIFAC output). Researched and added approximate GC constituent profiles for 9 blocked naturals (peppermint, spearmint, basil, roman camomile, armoise, bay laurel, tarragon, taget, copaiba) and expanded them into single-CAS constituents.
  • Compliance leak caught and closed: an intermediate pass simulated 23 rows without re-applying the banned substitution, leaving banned CAS in trainable rows. Fixed; final scan: 0 banned CAS in any trainable row.

Final empirical state: 5,678 / 5,708 trainable (99.47%). 177 banned- material rows recovered into the trainable set. 30 rows remain blocked β€” all natural oils (costus, basil variants, tarragon, camomile, copaiba, etc.) whose exact GC compositions are not in any local source; blocked rows are isolated (repair_status set), not trained. Remaining unblock = sourcing/authoring constituent profiles for those 14 naturals. Tests 111/111 green.