Buckets:

10.2 TB
4,599 files
Updated 7 days ago
Name
Size
test
tokenizer
train
README.md3.1 kB
xet
USAGE.md5.17 kB
xet
load_data.py2.3 kB
xet
manifest.jsonl1.66 MB
xet
release.json985 Bytes
xet
schema-observations.json3.61 kB
xet
README.md

Carbon-A packed training and evaluation data

Reusable HDF5 data for the GENERanno eukaryotic CDS annotator (Carbon-A). This bucket is currently private.

Where the data come from

The examples were constructed from paired RefSeq GBFF annotations and FASTA sequences, using GCF-accessioned assemblies from mammals, other vertebrates, invertebrates, plants, fungi, and protozoa. Only records at least 98,304 bp long were eligible. Domain-specific overlapping windows and small boundary offsets increase exposure to less abundant groups, particularly fungi and protozoa.

Each example has separate forward- and reverse-strand CDS targets. CDS intervals across isoforms are combined within each strand; transcript paths are not represented. Windows with no annotated CDS are retained. The saved files already contain the augmentation and 6-mer tokenization.

What is included

Location Contents
train/<division>/*.h5 4,561 original training shards; 10.16 TB
test/<genome>/<GCF-accession>.h5 Original 26-genome evaluation panel; 22.13 GB
tokenizer/ Tokenizer/config snapshot matching the model
load_data.py Minimal PyTorch dataset loader
manifest.jsonl, release.json File sizes, content hashes, and source provenance

Sizes are decimal. Mammals and other vertebrates share train/vertebrate/. The files use the original 16,384-token context, without mixing in shorter-context repacks or ablation subsets.

test/ preserves the evaluation panel used during training; it is not a newly constructed, untouched test split. Keep it separate from training. This packaging does not establish a new accession/species-disjointness audit.

Train with the packed files

Each row contains input_ids[16384] (uint16), label_plus[98304], and label_minus[98304] (int8). One token covers six bases.

Convert each stored label track with label != 0, because raw labels may be signed. Concatenate the full forward track followed by the reverse track, producing 196,608 targets. Do not add BOS/EOS tokens or reverse the stored tracks. The supplied loader handles this conversion and masks special-token label positions with -100.

After downloading selected shards and load_data.py:

from torch.utils.data import DataLoader
from load_data import PackedAnnotationDataset

data = PackedAnnotationDataset(["data/fungi.h5"])
batch = next(iter(DataLoader(data, batch_size=1, shuffle=True)))
# Move batch tensors to your training device, then use:
# outputs = model(**batch)
# outputs.loss.backward()

Use a compatible two-strand classification model with cross-entropy loss and ignore_index=-100. See USAGE.md in this bucket for installation, a one-shard download example, and context-length handling. Training HDF5 files contain tensors rather than per-window accession metadata; the separate metadata catalog supports provenance exploration.

Total size
10.2 TB
Files
4,599
Last updated
Oct 1
Pre-warmed CDN
US EU US EU

Contributors