--- license: apache-2.0 library_name: transformers tags: - dna - genomics - dna-language-model - mamba - moe pipeline_tag: text-generation --- # CENO-1B-base **CENO-1B-base** is the base pretraining checkpoint of the 1B **CENO** DNA foundation model — a causal language model over genomic sequence built on a Nemotron-H **Mamba / Attention / Mixture-of-Experts** hybrid backbone (no MSA inputs). It is part of the **CENO** DNA foundation model family. Model code, the VEP pipeline, and a generation demo live in the companion [CENO code repository](https://github.com/CladeTeam/CENO). This checkpoint is standalone-loadable with `trust_remote_code=True` — the model code is bundled here. ## Model details | | | |---|---| | Family | CENO (base) | | Training stage | Base pretraining (stage 2) | | Parameters | 1.3B (1302.4M) | | Precision | bfloat16 | | `model_type` | `ceno` | | Architecture class | `CENOForCausalLM` | | Auto-map (model) | `modeling_ceno.CENOForCausalLM` | | Auto-map (tokenizer) | `ceno_tokenizer.CENOCharLevelTokenizer` | ## Architecture | Property | Value | |---|---| | Hidden layers | 38 | | Hidden size | 1024 | | Attention heads | 16 | | Intermediate size | 4096 | | Experts (MoE) | 8 (top-2 per token) | | Vocabulary | 512 (byte / character-level) | The backbone is a Mamba / Attention / Mixture-of-Experts hybrid (Nemotron-H architecture). The tokenizer is character-level, mapping DNA bases to their ASCII byte codes. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer ckpt = "CladeTeam/CENO-1B-base" model = AutoModelForCausalLM.from_pretrained(ckpt, trust_remote_code=True) tokenizer = AutoTokenizer.from_pretrained(ckpt, trust_remote_code=True) ids = tokenizer.encode("ATCGATCG", return_tensors="pt") # out = model.generate(ids, max_new_tokens=128) # needs a CUDA GPU (Mamba kernels) ``` > The Mamba layers require CUDA kernels, so forward passes and generation need a GPU. > Config, tokenizer, and weight loading are CPU-safe. ## Intended use - **Base checkpoints (`CENO-*`)** — genomic-sequence generation and embedding extraction; downstream adaptation (fine-tuning, probing) for genomics tasks. - **MSA checkpoints (`CENO-P-*`)** — variant effect prediction (VEP) by scoring wild-type vs. variant sequences with delta log-likelihood. See the TraitGym VEP example in the [CENO code repository](https://github.com/CladeTeam/CENO). ## License Apache-2.0. The bundled model code is derived from NVIDIA's Nemotron-H Hugging Face implementation (Apache-2.0); the tokenizer is derived from the Arc Institute Evo2 `CharLevelTokenizer` (Apache-2.0). See the `LICENSE` and `NOTICE` files in this repository for full attribution.