Buckets:

cgeorgiaw's picture
|
download
raw
5.17 kB
# Carbon-A / GENERanno packed training data
Ready-to-load training and evaluation arrays for the [eukaryotic CDS annotator](https://huggingface.co/HuggingFaceBio/GENERanno-eukaryote-1.2b-cds-annotator).
## Origin
The data were constructed from paired RefSeq GBFF annotations and FASTA sequences from GCF-accessioned assemblies across mammals, other vertebrates, invertebrates, plants, fungi, and protozoa. Records shorter than **98,304 bp** were excluded. Eligible sequences were windowed with domain-specific overlapping strides and small boundary offsets, increasing the representation of fungi and protozoa. These windows are already tokenized; do not apply the augmentation again just to load them.
Targets mark CDS occupancy independently on the two strands. CDS intervals across annotated isoforms are combined within each strand; transcript paths are not represented. Windows without CDS are retained. The training interface uses **1 = coding, 0 = background**.
## Files
```text
train/{fungi,invertebrate,plant,protozoa,vertebrate}/*.h5
test/<evaluation-genome>/<GCF-accession>.h5
tokenizer/ # exact tokenizer/config snapshot
load_data.py # minimal PyTorch Dataset
manifest.jsonl # file sizes, source paths, and content hashes
release.json # pinned source revisions and release scope
```
There are **4,561 training files (10.16 TB)** and **26 evaluation files (22.13 GB)**; sizes are decimal. Mammals and other vertebrates share the `vertebrate` directory and remain distinguishable in filenames. This release uses the original **16,384-token** files, not the shorter-context repacks or ablation mixes.
`test/` preserves the original evaluation panel, which was used for evaluation during training. It is not a newly constructed, untouched test split. Keep it separate from training. This packaging does not establish an accession/species-disjointness audit; the training HDF5 arrays do not retain per-window accession metadata. The [separate metadata catalog](https://huggingface.co/datasets/HuggingFaceBio/Carbon-A-RefSeq-training) is available for provenance exploration.
Each HDF5 file contains:
| Array | Shape | Stored type |
| --- | --- | --- |
| `input_ids` | `[N, 16384]` | `uint16` |
| `label_plus` | `[N, 98304]` | `int8` |
| `label_minus` | `[N, 98304]` | `int8` |
One token represents six consecutive bases. **Raw labels can be signed and nonbinary; use `labels != 0`, not `labels > 0`.** The model expects all forward-strand base labels followed by all reverse-strand base labels, giving **196,608 labels per row**. Keep the reverse-strand track in its stored genomic coordinate order; do not reverse it or merge the two tracks.
## Download a small starting subset
Authenticate with a token that can read the private bucket:
```bash
pip install 'huggingface_hub==1.16.0' h5py numpy torch
hf auth login
```
Download one training shard (about 2.29 GB) and the loader:
```python
from pathlib import Path
from huggingface_hub import HfFileSystem
fs = HfFileSystem()
bucket = "buckets/HuggingFaceBio/Carbon-A-training-data"
path = "train/fungi/seq_fungi_part10_sub1.h5"
Path("data").mkdir(exist_ok=True)
fs.get_file(f"{bucket}/{path}", "data/fungi.h5")
fs.get_file(f"{bucket}/load_data.py", "load_data.py")
```
Use `manifest.jsonl` to select additional shards. For a full directory, `hf buckets sync hf://buckets/HuggingFaceBio/Carbon-A-training-data/train/fungi/ ./data/train/fungi/` downloads the fungi training collection; it is about 1 TB.
## Use in training
```python
from torch.utils.data import DataLoader
from load_data import PackedAnnotationDataset
dataset = PackedAnnotationDataset(["data/fungi.h5"])
loader = DataLoader(dataset, batch_size=1, shuffle=True, num_workers=0)
batch = next(iter(loader))
# input_ids, attention_mask: [1, 16384]
# labels: [1, 196608], with -100 at ignored positions
# After moving batch to your model's device:
# outputs = model(**batch)
# outputs.loss.backward()
```
The loader converts both label tracks to binary and masks the six base positions corresponding to `<oov>`, `<s>`, `</s>`, `<pad>`, and `<mask>` tokens with `-100`, matching the original training collator. These IDs are 0–4 in the bundled tokenizer. No BOS/EOS tokens should be added to the saved input IDs. If adapting to a shorter context, crop **each strand separately** at six times the token offsets, then concatenate; slicing the already-concatenated labels would misalign the strands.
Use a compatible two-strand annotation model with cross-entropy loss and `ignore_index=-100`. Full 98-kbp training requires suitable accelerator memory; choose batch size and distributed training settings for your hardware. This loader is a simple reuse example, not a reproduction of the full distributed training schedule.
The files are copied without changing their contents. `manifest.jsonl` records the original training bucket and the pinned evaluation repository revision; `release.json` identifies the model/tokenizer snapshot. RefSeq is the source of the biological sequences and annotations; this release does not assign a new license to those source records.

Xet Storage Details

Size:
5.17 kB
·
Xet hash:
08d4cef90f54df2cce98ce2a1f93e1b8dba53c8d4e480ce20082f83a4affd44a

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.