# scripts/data -- Data pipeline for the biopesticide-AI project This package builds the training CSV consumed by the downstream CNN/GAN model code. It produces a table where each row is a 21-nt siRNA sequence with: - the binarized efficacy label (`pest_label`; 1 if measured knockdown >= 0.7, else 0) -- this is the SCIENTIFICALLY CORRECT label, replacing the original merged code's wrong heuristic of labeling any 100-nt tile from a pest gene as `pest_label=1`, - 8 Reynolds 2004 siRNA efficacy features + a `reynolds_score`, - per-species off-target risk against the 6-species safety panel (apis_mellifera, bos_taurus, bos_indicus, gallus_gallus, danio_rerio, homo_sapiens). ## Files | File | Purpose | | ------------------------- | ------------------------------------------------------------------------------------ | | `__init__.py` | Empty; marks `scripts/data` as a Python package. | | `_legacy_imports.py` | Shim that re-exports `fasta_iter`, `KmerOffTargetIndex`, `tile_sequence`, etc. from the original `upload/merged_biopesticide_ai.py`. Copied verbatim (not imported) so the data pipeline does not depend on the soon-to-be-refactored `src/` layout. | | `reynolds_features.py` | `ReynoldsFeaturizer` class implementing the 8 Reynolds 2004 siRNA efficacy criteria. Pure Python + pandas, no torch. | | `download_sources.py` | Attempts to download siRecords, DRSC/TRiP, and genomeRNAi into `data/external/`. Verifies (but never re-downloads) the NCBI `rna.fna` files expected under `data//ncbi_dataset/data//`. All download failures are logged and skipped. | | `generate_synthetic.py` | Generates a small synthetic mini-dataset (50 pest transcripts, 1000 siRNAs split 500 high-efficacy / 500 low-efficacy, 50 safety transcripts across 5 species) so the pipeline can run end-to-end without 20GB of NCBI data. | | `build_training_csv.py` | The main orchestrator. Loads siRNAs, builds the off-target index, computes Reynolds features + off-target risk, binarizes the label, writes the final CSV. | ## How to run All commands are run from the project root (the directory containing the `bioai/` and `scripts/` packages) and use `python -m` so that intra-package imports resolve correctly. ### Synthetic path (recommended for development; no external data needed) ```bash # Step 1: generate the synthetic mini-dataset. python -m scripts.data.generate_synthetic # Step 2: build the final training CSV. python -m scripts.data.build_training_csv --source synthetic ``` Output: - `data/synthetic/pest_transcripts.fasta` - `data/synthetic/safety_transcripts.fasta` - `data/synthetic/sirna_training.csv` - `data/processed/training_data.csv` <-- consumed by the model code ### Real path (requires real siRNA CSV + NCBI rna.fna files on disk) ```bash # Step 1: attempt to fetch public RNAi databases (best-effort; will log # and skip on failure). This also verifies which NCBI files are # on disk. python -m scripts.data.download_sources # Step 2: build the final training CSV from real siRNA data + real NCBI # safety panel. Place your real siRNA CSV (with columns # `sirna_seq,target_gene,knockdown_pct`) at # `data/external/sirna_real.csv`, or pass --input . python -m scripts.data.build_training_csv --source real ``` Output: - `data/external/sirecords.csv` (only if the download succeeded) - `data/external/drsc_trip.csv` (only if the download succeeded) - `data/external/genomernai_phenotypes.txt` (only if the download succeeded) - `data/processed/training_data_real.csv` <-- consumed by the model code ## Output CSV schema `data/processed/training_data.csv` (synthetic) and `data/processed/training_data_real.csv` (real) have identical schemas: ``` sirna_seq, target_gene, knockdown_pct, pest_label, gc_content, no_repeats, at_pos19, a_pos3, t_pos10, ag_pos13, t_pos16, thermo_asymmetry, reynolds_score, offtarget_apis_mellifera, offtarget_bos_taurus, offtarget_bos_indicus, offtarget_gallus_gallus, offtarget_danio_rerio, offtarget_homo_sapiens ``` - `pest_label` = 1 if `knockdown_pct >= 0.7` else 0. - `reynolds_score` = integer 0-8, the count of Reynolds 2004 criteria met. - `offtarget_` = fraction of the siRNA's 21-mers (and their reverse complements) that appear in that species' transcriptome. 0.0 means no exact 21-mer overlap. ## Dependencies `pandas`, `numpy`, `pyyaml`, `biopython`, `tqdm`, `requests`. All CPU-only. No torch is required to run any of these scripts (the model code in `src/models/` will import the produced CSV). ## Reproducibility All synthetic generation seeds BOTH `numpy.random.default_rng(13)` and Python's `random.Random(13)` with the same seed (13). Re-running `generate_synthetic.py` produces byte-identical output. ## Known limitations - The synthetic siRNAs are random 21-mers, and the synthetic safety transcripts are also random. Exact 21-mer overlap between the two is effectively zero, so all `offtarget_` columns are 0.0 in the synthetic path. This is expected; the real-data path will produce meaningful off-target signal. - `KmerOffTargetIndex` uses exact 21-mer matching, which is strict. Real RNAi off-target effects often involve seed-region matches with mismatches; that's a future enhancement (out of scope for Task 3-B). - The public RNAi databases (siRecords, DRSC/TRiP, genomeRNAi) often have unstable download endpoints. `download_sources.py` tries several candidate URLs per source and logs+skips on failure; it will never block the pipeline. - `pest_label` is binarized from `knockdown_pct` at threshold 0.7. This is a simple heuristic; the spec asks for this exact threshold.