Buckets:
Carbon-A / GENERanno packed training data
Ready-to-load training and evaluation arrays for the eukaryotic CDS annotator.
Origin
The data were constructed from paired RefSeq GBFF annotations and FASTA sequences from GCF-accessioned assemblies across mammals, other vertebrates, invertebrates, plants, fungi, and protozoa. Records shorter than 98,304 bp were excluded. Eligible sequences were windowed with domain-specific overlapping strides and small boundary offsets, increasing the representation of fungi and protozoa. These windows are already tokenized; do not apply the augmentation again just to load them.
Targets mark CDS occupancy independently on the two strands. CDS intervals across annotated isoforms are combined within each strand; transcript paths are not represented. Windows without CDS are retained. The training interface uses 1 = coding, 0 = background.
Files
train/{fungi,invertebrate,plant,protozoa,vertebrate}/*.h5
test/<evaluation-genome>/<GCF-accession>.h5
tokenizer/ # exact tokenizer/config snapshot
load_data.py # minimal PyTorch Dataset
manifest.jsonl # file sizes, source paths, and content hashes
release.json # pinned source revisions and release scope
There are 4,561 training files (10.16 TB) and 26 evaluation files (22.13 GB); sizes are decimal. Mammals and other vertebrates share the vertebrate directory and remain distinguishable in filenames. This release uses the original 16,384-token files, not the shorter-context repacks or ablation mixes.
test/ preserves the original evaluation panel, which was used for evaluation during training. It is not a newly constructed, untouched test split. Keep it separate from training. This packaging does not establish an accession/species-disjointness audit; the training HDF5 arrays do not retain per-window accession metadata. The separate metadata catalog is available for provenance exploration.
Each HDF5 file contains:
| Array | Shape | Stored type |
|---|---|---|
input_ids |
[N, 16384] |
uint16 |
label_plus |
[N, 98304] |
int8 |
label_minus |
[N, 98304] |
int8 |
One token represents six consecutive bases. Raw labels can be signed and nonbinary; use labels != 0, not labels > 0. The model expects all forward-strand base labels followed by all reverse-strand base labels, giving 196,608 labels per row. Keep the reverse-strand track in its stored genomic coordinate order; do not reverse it or merge the two tracks.
Download a small starting subset
Authenticate with a token that can read the private bucket:
pip install 'huggingface_hub==1.16.0' h5py numpy torch
hf auth login
Download one training shard (about 2.29 GB) and the loader:
from pathlib import Path
from huggingface_hub import HfFileSystem
fs = HfFileSystem()
bucket = "buckets/HuggingFaceBio/Carbon-A-training-data"
path = "train/fungi/seq_fungi_part10_sub1.h5"
Path("data").mkdir(exist_ok=True)
fs.get_file(f"{bucket}/{path}", "data/fungi.h5")
fs.get_file(f"{bucket}/load_data.py", "load_data.py")
Use manifest.jsonl to select additional shards. For a full directory, hf buckets sync hf://buckets/HuggingFaceBio/Carbon-A-training-data/train/fungi/ ./data/train/fungi/ downloads the fungi training collection; it is about 1 TB.
Use in training
from torch.utils.data import DataLoader
from load_data import PackedAnnotationDataset
dataset = PackedAnnotationDataset(["data/fungi.h5"])
loader = DataLoader(dataset, batch_size=1, shuffle=True, num_workers=0)
batch = next(iter(loader))
# input_ids, attention_mask: [1, 16384]
# labels: [1, 196608], with -100 at ignored positions
# After moving batch to your model's device:
# outputs = model(**batch)
# outputs.loss.backward()
The loader converts both label tracks to binary and masks the six base positions corresponding to <oov>, <s>, </s>, <pad>, and <mask> tokens with -100, matching the original training collator. These IDs are 0–4 in the bundled tokenizer. No BOS/EOS tokens should be added to the saved input IDs. If adapting to a shorter context, crop each strand separately at six times the token offsets, then concatenate; slicing the already-concatenated labels would misalign the strands.
Use a compatible two-strand annotation model with cross-entropy loss and ignore_index=-100. Full 98-kbp training requires suitable accelerator memory; choose batch size and distributed training settings for your hardware. This loader is a simple reuse example, not a reproduction of the full distributed training schedule.
The files are copied without changing their contents. manifest.jsonl records the original training bucket and the pinned evaluation repository revision; release.json identifies the model/tokenizer snapshot. RefSeq is the source of the biological sequences and annotations; this release does not assign a new license to those source records.
Xet Storage Details
- Size:
- 5.17 kB
- Xet hash:
- 08d4cef90f54df2cce98ce2a1f93e1b8dba53c8d4e480ce20082f83a4affd44a
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.