Buckets:
| # Carbon-A / GENERanno packed training data | |
| Ready-to-load training and evaluation arrays for the [eukaryotic CDS annotator](https://huggingface.co/HuggingFaceBio/GENERanno-eukaryote-1.2b-cds-annotator). | |
| ## Origin | |
| The data were constructed from paired RefSeq GBFF annotations and FASTA sequences from GCF-accessioned assemblies across mammals, other vertebrates, invertebrates, plants, fungi, and protozoa. Records shorter than **98,304 bp** were excluded. Eligible sequences were windowed with domain-specific overlapping strides and small boundary offsets, increasing the representation of fungi and protozoa. These windows are already tokenized; do not apply the augmentation again just to load them. | |
| Targets mark CDS occupancy independently on the two strands. CDS intervals across annotated isoforms are combined within each strand; transcript paths are not represented. Windows without CDS are retained. The training interface uses **1 = coding, 0 = background**. | |
| ## Files | |
| ```text | |
| train/{fungi,invertebrate,plant,protozoa,vertebrate}/*.h5 | |
| test/<evaluation-genome>/<GCF-accession>.h5 | |
| tokenizer/ # exact tokenizer/config snapshot | |
| load_data.py # minimal PyTorch Dataset | |
| manifest.jsonl # file sizes, source paths, and content hashes | |
| release.json # pinned source revisions and release scope | |
| ``` | |
| There are **4,561 training files (10.16 TB)** and **26 evaluation files (22.13 GB)**; sizes are decimal. Mammals and other vertebrates share the `vertebrate` directory and remain distinguishable in filenames. This release uses the original **16,384-token** files, not the shorter-context repacks or ablation mixes. | |
| `test/` preserves the original evaluation panel, which was used for evaluation during training. It is not a newly constructed, untouched test split. Keep it separate from training. This packaging does not establish an accession/species-disjointness audit; the training HDF5 arrays do not retain per-window accession metadata. The [separate metadata catalog](https://huggingface.co/datasets/HuggingFaceBio/Carbon-A-RefSeq-training) is available for provenance exploration. | |
| Each HDF5 file contains: | |
| | Array | Shape | Stored type | | |
| | --- | --- | --- | | |
| | `input_ids` | `[N, 16384]` | `uint16` | | |
| | `label_plus` | `[N, 98304]` | `int8` | | |
| | `label_minus` | `[N, 98304]` | `int8` | | |
| One token represents six consecutive bases. **Raw labels can be signed and nonbinary; use `labels != 0`, not `labels > 0`.** The model expects all forward-strand base labels followed by all reverse-strand base labels, giving **196,608 labels per row**. Keep the reverse-strand track in its stored genomic coordinate order; do not reverse it or merge the two tracks. | |
| ## Download a small starting subset | |
| Authenticate with a token that can read the private bucket: | |
| ```bash | |
| pip install 'huggingface_hub==1.16.0' h5py numpy torch | |
| hf auth login | |
| ``` | |
| Download one training shard (about 2.29 GB) and the loader: | |
| ```python | |
| from pathlib import Path | |
| from huggingface_hub import HfFileSystem | |
| fs = HfFileSystem() | |
| bucket = "buckets/HuggingFaceBio/Carbon-A-training-data" | |
| path = "train/fungi/seq_fungi_part10_sub1.h5" | |
| Path("data").mkdir(exist_ok=True) | |
| fs.get_file(f"{bucket}/{path}", "data/fungi.h5") | |
| fs.get_file(f"{bucket}/load_data.py", "load_data.py") | |
| ``` | |
| Use `manifest.jsonl` to select additional shards. For a full directory, `hf buckets sync hf://buckets/HuggingFaceBio/Carbon-A-training-data/train/fungi/ ./data/train/fungi/` downloads the fungi training collection; it is about 1 TB. | |
| ## Use in training | |
| ```python | |
| from torch.utils.data import DataLoader | |
| from load_data import PackedAnnotationDataset | |
| dataset = PackedAnnotationDataset(["data/fungi.h5"]) | |
| loader = DataLoader(dataset, batch_size=1, shuffle=True, num_workers=0) | |
| batch = next(iter(loader)) | |
| # input_ids, attention_mask: [1, 16384] | |
| # labels: [1, 196608], with -100 at ignored positions | |
| # After moving batch to your model's device: | |
| # outputs = model(**batch) | |
| # outputs.loss.backward() | |
| ``` | |
| The loader converts both label tracks to binary and masks the six base positions corresponding to `<oov>`, `<s>`, `</s>`, `<pad>`, and `<mask>` tokens with `-100`, matching the original training collator. These IDs are 0–4 in the bundled tokenizer. No BOS/EOS tokens should be added to the saved input IDs. If adapting to a shorter context, crop **each strand separately** at six times the token offsets, then concatenate; slicing the already-concatenated labels would misalign the strands. | |
| Use a compatible two-strand annotation model with cross-entropy loss and `ignore_index=-100`. Full 98-kbp training requires suitable accelerator memory; choose batch size and distributed training settings for your hardware. This loader is a simple reuse example, not a reproduction of the full distributed training schedule. | |
| The files are copied without changing their contents. `manifest.jsonl` records the original training bucket and the pinned evaluation repository revision; `release.json` identifies the model/tokenizer snapshot. RefSeq is the source of the biological sequences and annotations; this release does not assign a new license to those source records. | |
Xet Storage Details
- Size:
- 5.17 kB
- Xet hash:
- 08d4cef90f54df2cce98ce2a1f93e1b8dba53c8d4e480ce20082f83a4affd44a
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.