Spaces:
Sleeping
scripts/data -- Data pipeline for the biopesticide-AI project
This package builds the training CSV consumed by the downstream CNN/GAN model code. It produces a table where each row is a 21-nt siRNA sequence with:
- the binarized efficacy label (
pest_label; 1 if measured knockdown >= 0.7, else 0) -- this is the SCIENTIFICALLY CORRECT label, replacing the original merged code's wrong heuristic of labeling any 100-nt tile from a pest gene aspest_label=1, - 8 Reynolds 2004 siRNA efficacy features + a
reynolds_score, - per-species off-target risk against the 6-species safety panel (apis_mellifera, bos_taurus, bos_indicus, gallus_gallus, danio_rerio, homo_sapiens).
Files
| File | Purpose |
|---|---|
__init__.py |
Empty; marks scripts/data as a Python package. |
_legacy_imports.py |
Shim that re-exports fasta_iter, KmerOffTargetIndex, tile_sequence, etc. from the original upload/merged_biopesticide_ai.py. Copied verbatim (not imported) so the data pipeline does not depend on the soon-to-be-refactored src/ layout. |
reynolds_features.py |
ReynoldsFeaturizer class implementing the 8 Reynolds 2004 siRNA efficacy criteria. Pure Python + pandas, no torch. |
download_sources.py |
Attempts to download siRecords, DRSC/TRiP, and genomeRNAi into data/external/. Verifies (but never re-downloads) the NCBI rna.fna files expected under data/<species>/ncbi_dataset/data/<assembly>/. All download failures are logged and skipped. |
generate_synthetic.py |
Generates a small synthetic mini-dataset (50 pest transcripts, 1000 siRNAs split 500 high-efficacy / 500 low-efficacy, 50 safety transcripts across 5 species) so the pipeline can run end-to-end without 20GB of NCBI data. |
build_training_csv.py |
The main orchestrator. Loads siRNAs, builds the off-target index, computes Reynolds features + off-target risk, binarizes the label, writes the final CSV. |
How to run
All commands are run from the project root (the directory containing the
bioai/ and scripts/ packages) and use python -m so that intra-package
imports resolve correctly.
Synthetic path (recommended for development; no external data needed)
# Step 1: generate the synthetic mini-dataset.
python -m scripts.data.generate_synthetic
# Step 2: build the final training CSV.
python -m scripts.data.build_training_csv --source synthetic
Output:
data/synthetic/pest_transcripts.fastadata/synthetic/safety_transcripts.fastadata/synthetic/sirna_training.csvdata/processed/training_data.csv<-- consumed by the model code
Real path (requires real siRNA CSV + NCBI rna.fna files on disk)
# Step 1: attempt to fetch public RNAi databases (best-effort; will log
# and skip on failure). This also verifies which NCBI files are
# on disk.
python -m scripts.data.download_sources
# Step 2: build the final training CSV from real siRNA data + real NCBI
# safety panel. Place your real siRNA CSV (with columns
# `sirna_seq,target_gene,knockdown_pct`) at
# `data/external/sirna_real.csv`, or pass --input <path>.
python -m scripts.data.build_training_csv --source real
Output:
data/external/sirecords.csv(only if the download succeeded)data/external/drsc_trip.csv(only if the download succeeded)data/external/genomernai_phenotypes.txt(only if the download succeeded)data/processed/training_data_real.csv<-- consumed by the model code
Output CSV schema
data/processed/training_data.csv (synthetic) and
data/processed/training_data_real.csv (real) have identical schemas:
sirna_seq, target_gene, knockdown_pct, pest_label,
gc_content, no_repeats, at_pos19, a_pos3, t_pos10, ag_pos13,
t_pos16, thermo_asymmetry, reynolds_score,
offtarget_apis_mellifera, offtarget_bos_taurus,
offtarget_bos_indicus, offtarget_gallus_gallus,
offtarget_danio_rerio, offtarget_homo_sapiens
pest_label= 1 ifknockdown_pct >= 0.7else 0.reynolds_score= integer 0-8, the count of Reynolds 2004 criteria met.offtarget_<species>= fraction of the siRNA's 21-mers (and their reverse complements) that appear in that species' transcriptome. 0.0 means no exact 21-mer overlap.
Dependencies
pandas, numpy, pyyaml, biopython, tqdm, requests. All CPU-only.
No torch is required to run any of these scripts (the model code in
src/models/ will import the produced CSV).
Reproducibility
All synthetic generation seeds BOTH numpy.random.default_rng(13) and
Python's random.Random(13) with the same seed (13). Re-running
generate_synthetic.py produces byte-identical output.
Known limitations
- The synthetic siRNAs are random 21-mers, and the synthetic safety
transcripts are also random. Exact 21-mer overlap between the two is
effectively zero, so all
offtarget_<species>columns are 0.0 in the synthetic path. This is expected; the real-data path will produce meaningful off-target signal. KmerOffTargetIndexuses exact 21-mer matching, which is strict. Real RNAi off-target effects often involve seed-region matches with mismatches; that's a future enhancement (out of scope for Task 3-B).- The public RNAi databases (siRecords, DRSC/TRiP, genomeRNAi) often
have unstable download endpoints.
download_sources.pytries several candidate URLs per source and logs+skips on failure; it will never block the pipeline. pest_labelis binarized fromknockdown_pctat threshold 0.7. This is a simple heuristic; the spec asks for this exact threshold.