File size: 12,179 Bytes
c87881a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 | # BGC retrieval rebuild
This directory is an isolated, leakage-free replacement for the legacy experiment.
It reads legacy inputs but never rewrites them. The primary development task is
within-species retrieval of held-out MIBiG-reference groups in the
*Streptomyces griseus* atlas. Curated within-species product labels are still
required for a publication-grade biological claim.
## Scientific invariants
1. Model inputs are restricted to frozen ESM-2 gene embeddings and relative gene
position. Alignment statistics and upstream product/activity predictions are
forbidden.
2. Product/reference groups are split before pairs or batches are constructed.
Test groups are excluded from pretraining, supervised training, tuning, and
checkpoint selection.
3. Atlas-to-MIBiG assignments are silver training labels. Curated *S. griseus*
BGC–product mappings are required for the primary biological result; broad
multi-genus benchmarking is secondary stress evidence only.
4. Ensemble weights are selected on validation once. Tied baseline scores use
expected metrics under uniform ordering within each tie group.
5. Every checkpoint and metric file records configuration, split, input, code,
and environment fingerprints.
## Layout and workflow
- `configs/main.yaml`: canonical one-GPU campaign configuration.
- `src/bgc_retrieval`: data contracts, model, training, evaluation, and audit CLI.
- `data/manifests`: tracked hashes, split assignments, and exclusion reports.
- `artifacts/<run_id>`: immutable run outputs; never overwrite an existing run.
From this directory, create a modern environment and install the package:
```bash
python -m venv .venv
.venv/bin/pip install -e '.[dev]'
bgc-rebuild audit --config configs/main.yaml
bgc-rebuild build-splits --config configs/main.yaml
bgc-rebuild train --config configs/main.yaml --stage phase1
bgc-rebuild train --config configs/main.yaml --stage phase2
bgc-rebuild evaluate --config configs/main.yaml
```
The legacy CSV has MIBiG-reference assignments but no curated product identity
table. `build-splits` therefore creates a **silver development split** by reference
ID and refuses to call it product-level. The primary development endpoint is
held-out, family-aggregated Recall@50 within the *S. griseus* atlas. The staged
multi-genus benchmark measures out-of-distribution correlation between BGC
embedding similarity and product-structure Tanimoto; it is not the paper's
species-scoped primary endpoint. Exact-product retrieval remains secondary because
most exact-product groups are singletons.
Prepare and run the secondary stress endpoint with:
~~~bash
python scripts/prepare_external_benchmark.py
python scripts/embed_external_proteins.py --resume
python scripts/evaluate_external.py \
--checkpoint artifacts/<training-run>/phase2_best.pt \
--run-id <external-evaluation-run>
~~~
`scripts/dgx_campaign.sh CAMPAIGN_TAG` runs the complete internal and external
campaign on a CUDA host. It reuses the host's CUDA-enabled PyTorch installation,
refuses to overwrite immutable run directories, and accepts
`EXTERNAL_BOOTSTRAP_SAMPLES` for a shorter smoke campaign. The default is 1,000.
## 2026-08-11 scope correction and residual redesign report
This section is the authoritative status for the rebuild and supersedes any earlier
wording that treats the multi-genus benchmark as the primary paper endpoint.
### Authoritative biological scope
The project is restricted to **within-species retrieval in _Streptomyces
griseus_**: given known members of a held-out compound/reference family, rank
other BGCs from the _S. griseus_ atlas. Generalization to other species or genera
is not a primary claim.
The legacy atlas contains 6,953 BGCs from 182 deduplicated genome accessions. Of
those BGCs, 5,325 have silver MIBiG-reference assignments used by the rebuild.
The frozen group-disjoint split contains:
| Split | BGCs | reference groups |
|---|---:|---:|
| train | 4,211 | 98 |
| validation | 676 | 21 |
| test | 438 | 21 |
Only 16 test groups contain more than the four required reference BGCs and
therefore contribute to the reported retrieval summaries. The atlas itself has a
`genome_id` column but no organism/species column. Its _S. griseus_ identity is
currently a dataset-provenance assertion, not a machine-verifiable row-level
contract. A taxonomy manifest for all 182 assemblies is required before
submission.
### Why the previous analysis went wrong
1. **The scientific task drifted.** The rebuild promoted the released
BGC-clustering benchmark to the primary endpoint even though the paper studies
only _S. griseus_. That external sequence set contains 1,343 BGCs spanning
hundreds of taxa: the unambiguous metadata contain 239 genera, and only 11
sequence records mention _S. griseus_. The resulting 1,270-BGC evaluation was
principally an out-of-distribution, multi-genus test. It is useful as a stress
test but does not answer the paper's within-species question.
2. **The labels do not constitute gold product identity.** The atlas
`compound_family` values were produced from MIBiG alignment evidence. The
rebuild correctly calls them silver `mibig_reference_id` groups, but the
legacy manuscript sometimes described them as known compound families or
exact products. The leakage-free _S. griseus_ result is therefore development
evidence and is marked `publication_eligible: false`.
3. **The original encoder destroyed a strong baseline representation.** SetNet
projected every 1,280-dimensional ESM gene vector into 256 dimensions, passed
it through two induced-attention blocks, and pooled to another 256-dimensional
vector. It had 2,272,512 trainable parameters and no raw-ESM residual path.
This was excessive for only 98 training groups and allowed optimization to
overwrite useful pretrained geometry.
4. **The contrastive target was mis-specified.** Supervised contrastive loss
treated every different silver group as an equally strong negative. Related
metabolites and noisy MIBiG assignments can therefore become false negatives.
The loss optimized discrete silver-group separation, not the biological notion
of alternative routes to the same or related chemistry.
5. **Checkpoint selection and the broad stress endpoint were misaligned.**
Checkpoints were selected by silver validation Recall@50, while the promoted
external endpoint was Spearman correlation with continuous product-structure
Tanimoto similarity. Across the five original SetNet seeds, better validation
recall did not reliably imply better external structural correlation.
6. **Legacy model selection overstated evidence.** Earlier work selected ensemble
weights on the evaluated data, defined "hard" cases using weakness of the Pfam
baseline itself, and treated many repeated retrieval draws as if they were
independent observations. The rebuild fixed these issues with validation-only
selection, expected tie-aware metrics, family aggregation, paired
family-level tests, and Holm correction.
7. **The effective test size is small.** Although the atlas contains thousands of
BGCs, uncertainty is governed by 16 eligible held-out groups, not by BGC-pair
counts or 100 repeated draws per group. This produces wide intervals and makes
small metric changes unconvincing.
The prior multi-genus five-seed result remains recorded as a secondary
out-of-distribution stress test: SetNet Spearman was
`0.088884 ± 0.023216`, raw ESM was `0.114198`, and BiG-SCAPE was
`0.261492`. Those numbers must not be used to accept or reject the
_S. griseus_-scoped paper.
### Residual Gene Weighting v1
The executed redesign uses only frozen ESM embeddings and relative gene position:
- a 397,057-parameter gene-importance gate, 5.7 times smaller than SetNet;
- zero-initialized gate logits, making the starting representation exactly raw
ESM mean pooling;
- weighted pooling in the original 1,280-dimensional ESM space instead of a
destructive 256-dimensional bottleneck;
- soft Pfam/family relational distillation so related non-identical groups are
not automatically hard negatives;
- a frozen raw-ESM residual score with validation candidates including
`alpha=0`;
- a Pfam hybrid grid including pure Pfam;
- checkpoint and mixture selection on validation only;
- five deterministic seeds plus a five-checkpoint score ensemble.
The residual construction guarantees:
```text
cosine(residual_i, residual_j)
= (1 - alpha) * cosine(raw_i, raw_j)
+ alpha * cosine(weighted_i, weighted_j)
```
Thus `alpha=0` reproduces raw ESM exactly.
### Executed DGX result
All 26 tests passed before execution and again after artifact retrieval using
the DGX environment's standard-library test runner. Training seeds 20260810–20260814 each
early-stopped after 13 epochs. Every seed selected epoch 0 as its best checkpoint;
best validation Recall@50 was approximately `0.4656–0.4658`, while later
training reduced it to `0.4068–0.4346`.
All five individual evaluations and the five-checkpoint ensemble independently
selected:
- learned residual weight: `alpha=0.0`;
- hybrid learned-branch weight: `alpha=0.0`;
- raw-ESM weight within the raw/Pfam hybrid: `beta=0.4`.
Validation therefore rejected the trained gene-weighting branch. The fixed
_S. griseus_ test results were:
| Method | Recall@50 | MRR | MAP | nDCG@50 |
|---|---:|---:|---:|---:|
| raw ESM mean | 0.792383 | 0.253128 | 0.724006 | 0.805809 |
| trained weighted ESM, five-seed mean | 0.792284 | 0.252970 | 0.723550 | 0.805465 |
| Pfam Jaccard max | 0.877924 | 0.305931 | 0.847764 | 0.903320 |
| validation-selected raw/Pfam hybrid | 0.878192 | 0.307223 | 0.849608 | 0.905321 |
The trained weighted representation was slightly worse than raw ESM. The
validation-selected residual correctly fell back to raw ESM. The raw/Pfam hybrid
improved Recall@50 over Pfam by only `0.000268`; the paired family-level
Holm-adjusted p-value was `0.498`. This is not a meaningful or statistically
supported improvement.
**Conclusion:** Residual Gene Weighting v1 successfully prevents degradation but
does not rescue the learned-method claim. The current evidence supports Pfam
Jaccard as the strongest _S. griseus_ method. Further model tuning on the same 16
test groups would be test-set overfitting.
### Reproducibility
Canonical configuration and entry points:
```bash
python scripts/train_residual.py \
--config configs/residual_griseus.yaml \
--run-id paper-v2-residual-seed-20260810
python scripts/evaluate_residual.py \
--config configs/residual_griseus.yaml \
--checkpoint artifacts/paper-v2-residual-seed-20260810/residual_best.pt \
--run-id paper-v2-residual-seed-20260810-evaluation
python scripts/analyze_residual_campaign.py \
--seed 20260810 --seed 20260811 --seed 20260812 \
--seed 20260813 --seed 20260814 \
--ensemble-run paper-v2-residual-five-seed-ensemble-evaluation \
--output-dir artifacts/paper-v2-residual-analysis
```
Primary result files:
- `artifacts/paper-v2-residual-analysis/aggregate_summary.csv`
- `artifacts/paper-v2-residual-analysis/paired_comparisons.csv`
- `artifacts/paper-v2-residual-analysis/analysis_metadata.json`
- `artifacts/paper-v2-residual-five-seed-ensemble-evaluation/summary.csv`
- `artifacts/paper-v2-residual-campaign.log`
- `artifacts/paper-v2-residual-sha256.txt` (58 files; manifest SHA-256
`2cee5df67735490d886b94e642375bb771785bb7547ab02ed90ad000f4e511ee`)
### Required next scientific step
Do not tune another architecture against the same test groups. First:
1. create a checked taxonomy manifest proving the 182 assemblies are
_S. griseus_;
2. curate or obtain gold BGC–product labels within those genomes;
3. version a new untouched family-level holdout or use fully nested group
cross-validation;
4. pre-register the primary metric and baseline comparisons;
5. only then train a chemistry-aligned residual model.
### Documentation rule
After every future request that changes code, data, experiments, results, scope,
or scientific interpretation, update this README in the same request with a dated
summary and artifact paths.
|