flvcko's picture
Biopesticide-AI: AMD Hackathon Unicorn Track submission
914512c
|
Raw
History Blame Contribute Delete
5.84 kB

scripts/data -- Data pipeline for the biopesticide-AI project

This package builds the training CSV consumed by the downstream CNN/GAN model code. It produces a table where each row is a 21-nt siRNA sequence with:

  • the binarized efficacy label (pest_label; 1 if measured knockdown >= 0.7, else 0) -- this is the SCIENTIFICALLY CORRECT label, replacing the original merged code's wrong heuristic of labeling any 100-nt tile from a pest gene as pest_label=1,
  • 8 Reynolds 2004 siRNA efficacy features + a reynolds_score,
  • per-species off-target risk against the 6-species safety panel (apis_mellifera, bos_taurus, bos_indicus, gallus_gallus, danio_rerio, homo_sapiens).

Files

File Purpose
__init__.py Empty; marks scripts/data as a Python package.
_legacy_imports.py Shim that re-exports fasta_iter, KmerOffTargetIndex, tile_sequence, etc. from the original upload/merged_biopesticide_ai.py. Copied verbatim (not imported) so the data pipeline does not depend on the soon-to-be-refactored src/ layout.
reynolds_features.py ReynoldsFeaturizer class implementing the 8 Reynolds 2004 siRNA efficacy criteria. Pure Python + pandas, no torch.
download_sources.py Attempts to download siRecords, DRSC/TRiP, and genomeRNAi into data/external/. Verifies (but never re-downloads) the NCBI rna.fna files expected under data/<species>/ncbi_dataset/data/<assembly>/. All download failures are logged and skipped.
generate_synthetic.py Generates a small synthetic mini-dataset (50 pest transcripts, 1000 siRNAs split 500 high-efficacy / 500 low-efficacy, 50 safety transcripts across 5 species) so the pipeline can run end-to-end without 20GB of NCBI data.
build_training_csv.py The main orchestrator. Loads siRNAs, builds the off-target index, computes Reynolds features + off-target risk, binarizes the label, writes the final CSV.

How to run

All commands are run from the project root (the directory containing the bioai/ and scripts/ packages) and use python -m so that intra-package imports resolve correctly.

Synthetic path (recommended for development; no external data needed)

# Step 1: generate the synthetic mini-dataset.
python -m scripts.data.generate_synthetic

# Step 2: build the final training CSV.
python -m scripts.data.build_training_csv --source synthetic

Output:

  • data/synthetic/pest_transcripts.fasta
  • data/synthetic/safety_transcripts.fasta
  • data/synthetic/sirna_training.csv
  • data/processed/training_data.csv <-- consumed by the model code

Real path (requires real siRNA CSV + NCBI rna.fna files on disk)

# Step 1: attempt to fetch public RNAi databases (best-effort; will log
#         and skip on failure). This also verifies which NCBI files are
#         on disk.
python -m scripts.data.download_sources

# Step 2: build the final training CSV from real siRNA data + real NCBI
#         safety panel. Place your real siRNA CSV (with columns
#         `sirna_seq,target_gene,knockdown_pct`) at
#         `data/external/sirna_real.csv`, or pass --input <path>.
python -m scripts.data.build_training_csv --source real

Output:

  • data/external/sirecords.csv (only if the download succeeded)
  • data/external/drsc_trip.csv (only if the download succeeded)
  • data/external/genomernai_phenotypes.txt (only if the download succeeded)
  • data/processed/training_data_real.csv <-- consumed by the model code

Output CSV schema

data/processed/training_data.csv (synthetic) and data/processed/training_data_real.csv (real) have identical schemas:

sirna_seq, target_gene, knockdown_pct, pest_label,
gc_content, no_repeats, at_pos19, a_pos3, t_pos10, ag_pos13,
t_pos16, thermo_asymmetry, reynolds_score,
offtarget_apis_mellifera, offtarget_bos_taurus,
offtarget_bos_indicus, offtarget_gallus_gallus,
offtarget_danio_rerio, offtarget_homo_sapiens
  • pest_label = 1 if knockdown_pct >= 0.7 else 0.
  • reynolds_score = integer 0-8, the count of Reynolds 2004 criteria met.
  • offtarget_<species> = fraction of the siRNA's 21-mers (and their reverse complements) that appear in that species' transcriptome. 0.0 means no exact 21-mer overlap.

Dependencies

pandas, numpy, pyyaml, biopython, tqdm, requests. All CPU-only. No torch is required to run any of these scripts (the model code in src/models/ will import the produced CSV).

Reproducibility

All synthetic generation seeds BOTH numpy.random.default_rng(13) and Python's random.Random(13) with the same seed (13). Re-running generate_synthetic.py produces byte-identical output.

Known limitations

  • The synthetic siRNAs are random 21-mers, and the synthetic safety transcripts are also random. Exact 21-mer overlap between the two is effectively zero, so all offtarget_<species> columns are 0.0 in the synthetic path. This is expected; the real-data path will produce meaningful off-target signal.
  • KmerOffTargetIndex uses exact 21-mer matching, which is strict. Real RNAi off-target effects often involve seed-region matches with mismatches; that's a future enhancement (out of scope for Task 3-B).
  • The public RNAi databases (siRecords, DRSC/TRiP, genomeRNAi) often have unstable download endpoints. download_sources.py tries several candidate URLs per source and logs+skips on failure; it will never block the pipeline.
  • pest_label is binarized from knockdown_pct at threshold 0.7. This is a simple heuristic; the spec asks for this exact threshold.