Buckets:

cgeorgiaw's picture
|
download
raw
5.17 kB

Carbon-A / GENERanno packed training data

Ready-to-load training and evaluation arrays for the eukaryotic CDS annotator.

Origin

The data were constructed from paired RefSeq GBFF annotations and FASTA sequences from GCF-accessioned assemblies across mammals, other vertebrates, invertebrates, plants, fungi, and protozoa. Records shorter than 98,304 bp were excluded. Eligible sequences were windowed with domain-specific overlapping strides and small boundary offsets, increasing the representation of fungi and protozoa. These windows are already tokenized; do not apply the augmentation again just to load them.

Targets mark CDS occupancy independently on the two strands. CDS intervals across annotated isoforms are combined within each strand; transcript paths are not represented. Windows without CDS are retained. The training interface uses 1 = coding, 0 = background.

Files

train/{fungi,invertebrate,plant,protozoa,vertebrate}/*.h5
test/<evaluation-genome>/<GCF-accession>.h5
tokenizer/                  # exact tokenizer/config snapshot
load_data.py                # minimal PyTorch Dataset
manifest.jsonl              # file sizes, source paths, and content hashes
release.json                # pinned source revisions and release scope

There are 4,561 training files (10.16 TB) and 26 evaluation files (22.13 GB); sizes are decimal. Mammals and other vertebrates share the vertebrate directory and remain distinguishable in filenames. This release uses the original 16,384-token files, not the shorter-context repacks or ablation mixes.

test/ preserves the original evaluation panel, which was used for evaluation during training. It is not a newly constructed, untouched test split. Keep it separate from training. This packaging does not establish an accession/species-disjointness audit; the training HDF5 arrays do not retain per-window accession metadata. The separate metadata catalog is available for provenance exploration.

Each HDF5 file contains:

Array Shape Stored type
input_ids [N, 16384] uint16
label_plus [N, 98304] int8
label_minus [N, 98304] int8

One token represents six consecutive bases. Raw labels can be signed and nonbinary; use labels != 0, not labels > 0. The model expects all forward-strand base labels followed by all reverse-strand base labels, giving 196,608 labels per row. Keep the reverse-strand track in its stored genomic coordinate order; do not reverse it or merge the two tracks.

Download a small starting subset

Authenticate with a token that can read the private bucket:

pip install 'huggingface_hub==1.16.0' h5py numpy torch
hf auth login

Download one training shard (about 2.29 GB) and the loader:

from pathlib import Path
from huggingface_hub import HfFileSystem

fs = HfFileSystem()
bucket = "buckets/HuggingFaceBio/Carbon-A-training-data"
path = "train/fungi/seq_fungi_part10_sub1.h5"
Path("data").mkdir(exist_ok=True)
fs.get_file(f"{bucket}/{path}", "data/fungi.h5")
fs.get_file(f"{bucket}/load_data.py", "load_data.py")

Use manifest.jsonl to select additional shards. For a full directory, hf buckets sync hf://buckets/HuggingFaceBio/Carbon-A-training-data/train/fungi/ ./data/train/fungi/ downloads the fungi training collection; it is about 1 TB.

Use in training

from torch.utils.data import DataLoader
from load_data import PackedAnnotationDataset

dataset = PackedAnnotationDataset(["data/fungi.h5"])
loader = DataLoader(dataset, batch_size=1, shuffle=True, num_workers=0)
batch = next(iter(loader))
# input_ids, attention_mask: [1, 16384]
# labels: [1, 196608], with -100 at ignored positions
# After moving batch to your model's device:
# outputs = model(**batch)
# outputs.loss.backward()

The loader converts both label tracks to binary and masks the six base positions corresponding to <oov>, <s>, </s>, <pad>, and <mask> tokens with -100, matching the original training collator. These IDs are 0–4 in the bundled tokenizer. No BOS/EOS tokens should be added to the saved input IDs. If adapting to a shorter context, crop each strand separately at six times the token offsets, then concatenate; slicing the already-concatenated labels would misalign the strands.

Use a compatible two-strand annotation model with cross-entropy loss and ignore_index=-100. Full 98-kbp training requires suitable accelerator memory; choose batch size and distributed training settings for your hardware. This loader is a simple reuse example, not a reproduction of the full distributed training schedule.

The files are copied without changing their contents. manifest.jsonl records the original training bucket and the pinned evaluation repository revision; release.json identifies the model/tokenizer snapshot. RefSeq is the source of the biological sequences and annotations; this release does not assign a new license to those source records.

Xet Storage Details

Size:
5.17 kB
·
Xet hash:
08d4cef90f54df2cce98ce2a1f93e1b8dba53c8d4e480ce20082f83a4affd44a

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.