Buckets:

cgeorgiaw's picture
|
download
raw
3.1 kB
# Carbon-A packed training and evaluation data
Reusable HDF5 data for the [GENERanno eukaryotic CDS annotator (Carbon-A)](https://huggingface.co/HuggingFaceBio/GENERanno-eukaryote-1.2b-cds-annotator). This bucket is currently private.
## Where the data come from
The examples were constructed from paired **RefSeq GBFF annotations and FASTA sequences**, using GCF-accessioned assemblies from mammals, other vertebrates, invertebrates, plants, fungi, and protozoa. Only records at least **98,304 bp** long were eligible. Domain-specific overlapping windows and small boundary offsets increase exposure to less abundant groups, particularly fungi and protozoa.
Each example has separate forward- and reverse-strand CDS targets. CDS intervals across isoforms are combined within each strand; transcript paths are not represented. Windows with no annotated CDS are retained. The saved files already contain the augmentation and 6-mer tokenization.
## What is included
| Location | Contents |
| --- | --- |
| `train/<division>/*.h5` | 4,561 original training shards; 10.16 TB |
| `test/<genome>/<GCF-accession>.h5` | Original 26-genome evaluation panel; 22.13 GB |
| `tokenizer/` | Tokenizer/config snapshot matching the model |
| `load_data.py` | Minimal PyTorch dataset loader |
| `manifest.jsonl`, `release.json` | File sizes, content hashes, and source provenance |
Sizes are decimal. Mammals and other vertebrates share `train/vertebrate/`. The files use the original **16,384-token context**, without mixing in shorter-context repacks or ablation subsets.
`test/` preserves the evaluation panel used during training; it is not a newly constructed, untouched test split. Keep it separate from training. This packaging does not establish a new accession/species-disjointness audit.
## Train with the packed files
Each row contains `input_ids[16384]` (`uint16`), `label_plus[98304]`, and `label_minus[98304]` (`int8`). One token covers six bases.
**Convert each stored label track with `label != 0`**, because raw labels may be signed. Concatenate the full forward track followed by the reverse track, producing 196,608 targets. Do not add BOS/EOS tokens or reverse the stored tracks. The supplied loader handles this conversion and masks special-token label positions with `-100`.
After downloading selected shards and `load_data.py`:
```python
from torch.utils.data import DataLoader
from load_data import PackedAnnotationDataset
data = PackedAnnotationDataset(["data/fungi.h5"])
batch = next(iter(DataLoader(data, batch_size=1, shuffle=True)))
# Move batch tensors to your training device, then use:
# outputs = model(**batch)
# outputs.loss.backward()
```
Use a compatible two-strand classification model with cross-entropy loss and `ignore_index=-100`. See **USAGE.md** in this bucket for installation, a one-shard download example, and context-length handling. Training HDF5 files contain tensors rather than per-window accession metadata; the [separate metadata catalog](https://huggingface.co/datasets/HuggingFaceBio/Carbon-A-RefSeq-training) supports provenance exploration.

Xet Storage Details

Size:
3.1 kB
·
Xet hash:
ae55638a99b082c44be12f66b578fc3b37befd07764610f6841d68ee010c541c

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.