Buckets:
| # Carbon-A packed training and evaluation data | |
| Reusable HDF5 data for the [GENERanno eukaryotic CDS annotator (Carbon-A)](https://huggingface.co/HuggingFaceBio/GENERanno-eukaryote-1.2b-cds-annotator). This bucket is currently private. | |
| ## Where the data come from | |
| The examples were constructed from paired **RefSeq GBFF annotations and FASTA sequences**, using GCF-accessioned assemblies from mammals, other vertebrates, invertebrates, plants, fungi, and protozoa. Only records at least **98,304 bp** long were eligible. Domain-specific overlapping windows and small boundary offsets increase exposure to less abundant groups, particularly fungi and protozoa. | |
| Each example has separate forward- and reverse-strand CDS targets. CDS intervals across isoforms are combined within each strand; transcript paths are not represented. Windows with no annotated CDS are retained. The saved files already contain the augmentation and 6-mer tokenization. | |
| ## What is included | |
| | Location | Contents | | |
| | --- | --- | | |
| | `train/<division>/*.h5` | 4,561 original training shards; 10.16 TB | | |
| | `test/<genome>/<GCF-accession>.h5` | Original 26-genome evaluation panel; 22.13 GB | | |
| | `tokenizer/` | Tokenizer/config snapshot matching the model | | |
| | `load_data.py` | Minimal PyTorch dataset loader | | |
| | `manifest.jsonl`, `release.json` | File sizes, content hashes, and source provenance | | |
| Sizes are decimal. Mammals and other vertebrates share `train/vertebrate/`. The files use the original **16,384-token context**, without mixing in shorter-context repacks or ablation subsets. | |
| `test/` preserves the evaluation panel used during training; it is not a newly constructed, untouched test split. Keep it separate from training. This packaging does not establish a new accession/species-disjointness audit. | |
| ## Train with the packed files | |
| Each row contains `input_ids[16384]` (`uint16`), `label_plus[98304]`, and `label_minus[98304]` (`int8`). One token covers six bases. | |
| **Convert each stored label track with `label != 0`**, because raw labels may be signed. Concatenate the full forward track followed by the reverse track, producing 196,608 targets. Do not add BOS/EOS tokens or reverse the stored tracks. The supplied loader handles this conversion and masks special-token label positions with `-100`. | |
| After downloading selected shards and `load_data.py`: | |
| ```python | |
| from torch.utils.data import DataLoader | |
| from load_data import PackedAnnotationDataset | |
| data = PackedAnnotationDataset(["data/fungi.h5"]) | |
| batch = next(iter(DataLoader(data, batch_size=1, shuffle=True))) | |
| # Move batch tensors to your training device, then use: | |
| # outputs = model(**batch) | |
| # outputs.loss.backward() | |
| ``` | |
| Use a compatible two-strand classification model with cross-entropy loss and `ignore_index=-100`. See **USAGE.md** in this bucket for installation, a one-shard download example, and context-length handling. Training HDF5 files contain tensors rather than per-window accession metadata; the [separate metadata catalog](https://huggingface.co/datasets/HuggingFaceBio/Carbon-A-RefSeq-training) supports provenance exploration. | |
Xet Storage Details
- Size:
- 3.1 kB
- Xet hash:
- ae55638a99b082c44be12f66b578fc3b37befd07764610f6841d68ee010c541c
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.