Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| test | 26 items | ||
| tokenizer | 6 items | ||
| train | 4,561 items | ||
| README.md | 3.1 kB xet | ae55638a | |
| USAGE.md | 5.17 kB xet | 08d4cef9 | |
| load_data.py | 2.3 kB xet | 8a59076d | |
| manifest.jsonl | 1.66 MB xet | baacb675 | |
| release.json | 985 Bytes xet | f959882c | |
| schema-observations.json | 3.61 kB xet | a3801069 |
Carbon-A packed training and evaluation data
Reusable HDF5 data for the GENERanno eukaryotic CDS annotator (Carbon-A). This bucket is currently private.
Where the data come from
The examples were constructed from paired RefSeq GBFF annotations and FASTA sequences, using GCF-accessioned assemblies from mammals, other vertebrates, invertebrates, plants, fungi, and protozoa. Only records at least 98,304 bp long were eligible. Domain-specific overlapping windows and small boundary offsets increase exposure to less abundant groups, particularly fungi and protozoa.
Each example has separate forward- and reverse-strand CDS targets. CDS intervals across isoforms are combined within each strand; transcript paths are not represented. Windows with no annotated CDS are retained. The saved files already contain the augmentation and 6-mer tokenization.
What is included
| Location | Contents |
|---|---|
train/<division>/*.h5 |
4,561 original training shards; 10.16 TB |
test/<genome>/<GCF-accession>.h5 |
Original 26-genome evaluation panel; 22.13 GB |
tokenizer/ |
Tokenizer/config snapshot matching the model |
load_data.py |
Minimal PyTorch dataset loader |
manifest.jsonl, release.json |
File sizes, content hashes, and source provenance |
Sizes are decimal. Mammals and other vertebrates share train/vertebrate/. The files use the original 16,384-token context, without mixing in shorter-context repacks or ablation subsets.
test/ preserves the evaluation panel used during training; it is not a newly constructed, untouched test split. Keep it separate from training. This packaging does not establish a new accession/species-disjointness audit.
Train with the packed files
Each row contains input_ids[16384] (uint16), label_plus[98304], and label_minus[98304] (int8). One token covers six bases.
Convert each stored label track with label != 0, because raw labels may be signed. Concatenate the full forward track followed by the reverse track, producing 196,608 targets. Do not add BOS/EOS tokens or reverse the stored tracks. The supplied loader handles this conversion and masks special-token label positions with -100.
After downloading selected shards and load_data.py:
from torch.utils.data import DataLoader
from load_data import PackedAnnotationDataset
data = PackedAnnotationDataset(["data/fungi.h5"])
batch = next(iter(DataLoader(data, batch_size=1, shuffle=True)))
# Move batch tensors to your training device, then use:
# outputs = model(**batch)
# outputs.loss.backward()
Use a compatible two-strand classification model with cross-entropy loss and ignore_index=-100. See USAGE.md in this bucket for installation, a one-shard download example, and context-length handling. Training HDF5 files contain tensors rather than per-window accession metadata; the separate metadata catalog supports provenance exploration.
- Total size
- 10.2 TB
- Files
- 4,599
- Last updated
- Oct 1
- Pre-warmed CDN
- US EU US EU