Carbon-A-1.2B

A DNA annotation model from the Carbon family.

Carbon-A predicts protein-coding sequence (CDS) at single-base resolution in eukaryotic genomes. It produces separate forward- and reverse-strand probabilities, preserving coding regions that overlap on opposite strands.

Facts

  • 1.2B-parameter DNA annotation model with two strand-specific classification heads.
  • Tokenizer: non-overlapping 6-mers; each DNA token represents six bases.
  • Sequence length: 16,384 tokens (98,304 bp).
  • Inputs: genomic DNA. The examples support GenBank, FASTA, and packed HDF5 files.
  • Outputs: per-base probabilities for non-coding and CDS on each strand; binary calls use a configurable threshold, default 0.5.
  • Training data: RefSeq GCF assemblies from mammals, other vertebrates, invertebrates, plants, fungi, and protozoa.
  • Transformers interface: AutoModelForTokenClassification with trust_remote_code=True; Transformers 4.56 or later, including 5.x.
  • Inference: PyTorch SDPA, with BF16 on supported GPUs. A separate FlashAttention installation is not required.

How to use

Use a Python 3.10+ environment with CUDA-enabled PyTorch. Authenticate with your Hugging Face account, download the example scripts, and install their requirements:

pip install -U "huggingface_hub>=0.34"
hf auth login
hf download HuggingFaceBio/Carbon-A-1.2B --exclude model.safetensors --local-dir Carbon-A-1.2B
cd Carbon-A-1.2B
pip install -r requirements.txt

The scripts load the weights through from_pretrained, which downloads them (4.6 GB) to the Hugging Face cache on first use, so they are excluded above.

Full-length 98,304-base windows rely on PyTorch's fused SDPA kernels (FlashAttention or memory-efficient attention), which recent NVIDIA GPUs provide. Without them, PyTorch falls back to a kernel that materializes 16,384 × 16,384 attention matrices and needs more than 32 GB of additional GPU memory.

Use from Python

import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer

model_id = "HuggingFaceBio/Carbon-A-1.2B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForTokenClassification.from_pretrained(
    model_id, trust_remote_code=True, dtype=torch.bfloat16, attn_implementation="sdpa"
).to("cuda").eval()

sequence = "ATGTCTGCTACACAAAAGCAACACTAA"  # up to 98,304 bases
sequence = sequence.upper()
sequence += "N" * (-len(sequence) % 6)  # complete the final 6-mer
inputs = tokenizer(sequence, add_special_tokens=False, return_tensors="pt").to("cuda")
with torch.inference_mode():
    logits = model(**inputs).logits  # (1, 2 * len(sequence), 2)
cds_plus, cds_minus = logits.float().softmax(-1)[0, :, 1].reshape(2, -1).cpu()

Logits are per base: all forward-strand positions, then all reverse-strand positions, both in input coordinates. Positions covered by a 6-mer containing N or another ambiguous base are not meaningful; the example scripts exclude them. Pass attention_mask, which the tokenizer returns, because the model requires it. Uppercase the sequence and use add_special_tokens=False: the tokenizer adds a BOS token by default, and lowercase (soft-masked) bases are not grouped into 6-mers, which shifts every later position.

Annotate a GenBank or FASTA file

The included infer_genbank.py script reads the DNA sequence from each record and writes per-base predictions:

# GenBank flat file; gzip compression is supported.
python infer_genbank.py genome.gbff.gz --out-dir predictions

# FASTA input.
python infer_genbank.py genome.fna.gz --format fasta --out-dir predictions-fasta

Long records are processed in 98,304-base windows with 50% overlap. Probabilities are averaged where windows overlap. Existing GenBank annotations are not model inputs.

Each record produces an .npz file named with its accession and record name, plus an entry in summary.json. To read the predictions:

import json
from pathlib import Path
import numpy as np

out = Path("predictions")
record = json.loads((out / "summary.json").read_text())["records"][0]
with np.load(out / record["file"]) as pred:
    plus = pred["cds_probability_plus"]
    minus = pred["cds_probability_minus"]
    coding = pred["cds_label_any"]  # 1 if either strand passes the threshold
    valid = pred["valid"]

Array index i corresponds to record base i + 1 on both strands. Unscored positions have probability NaN, call -1, and valid=False.

Run on packed training or evaluation data

The included infer_packed.py script can download a small evaluation file and predict one full-length row:

python infer_packed.py --num-rows 1 --out-dir predictions-packed

Or use a local shard:

python infer_packed.py --h5 /path/to/shard.h5 --num-rows 2 --out-dir predictions-local

Both scripts accept --threshold and --revision <model-commit>. See the inference guide for output fields, tokenization, coordinate handling, and reproducible-run details.

Training data

Examples were constructed from paired RefSeq GBFF annotations and FASTA sequences, using GCF-accessioned assemblies. Records at least 98,304 bp long were eligible. Domain-specific sliding windows increase exposure to less abundant groups, particularly fungi and protozoa.

CDS intervals are unioned across isoforms independently on each strand. Windows without annotated CDS are retained, exposing the model to both coding and non-coding genomic context.

The packed-data bucket contains training shards, the original evaluation panel, a PyTorch loader, and reuse instructions. Stored labels are converted to binary with label != 0; the loader handles signed values and special-token masking.

Evaluation

The inference examples have passed functional checks on an NVIDIA H100, using BF16, PyTorch 2.7.1, and Transformers 4.56.0:

  • Two full-length packed evaluation windows, including strand outputs and thresholded calls.
  • The complete 230,218-base yeast chromosome I record, processed in four overlapping windows.
  • Identical predictions when the same sequence is supplied as GenBank or FASTA.
  • Independent checks of overlap averaging, coordinates, ambiguous bases, and short records.

These checks verify the inference code. They do not measure agreement with reference CDS annotations. Details are recorded in the packed-data check and the GenBank check.

Limitations

  • Predictions represent CDS occupancy. Transcript paths, alternative isoforms, and complete gene structures are not resolved.
  • The model is intended for eukaryotic sequence. Bacteria and viruses are outside its training scope.
  • The GenBank/FASTA example excludes any 6-mer containing an ambiguous nucleotide, including an incomplete final 6-mer padded with N.
  • Shorter records can be processed, but training used full-length 98,304-base windows.
  • The supplied evaluation panel was used during training; it is not a newly held-out test set.
  • Full-length inference on CPU is slow. The file example holds per-base arrays for one record in memory, so large chromosomes need more host RAM.

License

MIT. Apache-2.0 notices in the model code are retained.

Downloads last month
9
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including HuggingFaceBio/Carbon-A-1.2B

Article mentioning HuggingFaceBio/Carbon-A-1.2B