cuio
/

File size: 2,725 Bytes
4c6fb34
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
---
license: apache-2.0
library_name: transformers
tags:
  - dna
  - genomics
  - dna-language-model
  - mamba
  - moe
pipeline_tag: text-generation
---

# CENO-300M-base

**CENO-300M-base** is the base pretraining checkpoint of the 300M **CENO** DNA foundation model — a causal language model over genomic sequence built on a
Nemotron-H **Mamba / Attention / Mixture-of-Experts** hybrid backbone (no MSA inputs).

It is part of the **CENO** DNA foundation model family. Model code, the VEP pipeline, and a
generation demo live in the companion [CENO code repository](https://github.com/CladeTeam/CENO).
This checkpoint is standalone-loadable with `trust_remote_code=True` — the model code is
bundled here.

## Model details

|  |  |
|---|---|
| Family | CENO (base) |
| Training stage | Base pretraining (stage 2) |
| Parameters | 307.5M |
| Precision | bfloat16 |
| `model_type` | `ceno` |
| Architecture class | `CENOForCausalLM` |
| Auto-map (model) | `modeling_ceno.CENOForCausalLM` |
| Auto-map (tokenizer) | `ceno_tokenizer.CENOCharLevelTokenizer` |

## Architecture

| Property | Value |
|---|---|
| Hidden layers | 9 |
| Hidden size | 1024 |
| Attention heads | 8 |
| Intermediate size | 4096 |
| Experts (MoE) | 8 (top-2 per token) |
| Vocabulary | 512 (byte / character-level) |

The backbone is a Mamba / Attention / Mixture-of-Experts hybrid (Nemotron-H
architecture). The tokenizer is character-level, mapping DNA bases to their ASCII
byte codes.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

ckpt = "CladeTeam/CENO-300M-base"
model = AutoModelForCausalLM.from_pretrained(ckpt, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(ckpt, trust_remote_code=True)

ids = tokenizer.encode("ATCGATCG", return_tensors="pt")
# out = model.generate(ids, max_new_tokens=128)   # needs a CUDA GPU (Mamba kernels)
```

> The Mamba layers require CUDA kernels, so forward passes and generation need a GPU.
> Config, tokenizer, and weight loading are CPU-safe.

## Intended use

- **Base checkpoints (`CENO-*`)** — genomic-sequence generation and embedding extraction;
  downstream adaptation (fine-tuning, probing) for genomics tasks.
- **MSA checkpoints (`CENO-P-*`)** — variant effect prediction (VEP) by scoring wild-type
  vs. variant sequences with delta log-likelihood. See the TraitGym VEP example in the
  [CENO code repository](https://github.com/CladeTeam/CENO).

## License

Apache-2.0. The bundled model code is derived from NVIDIA's Nemotron-H Hugging Face
implementation (Apache-2.0); the tokenizer is derived from the Arc Institute Evo2
`CharLevelTokenizer` (Apache-2.0). See the `LICENSE` and `NOTICE` files in this repository
for full attribution.