How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("feature-extraction", model="Tim419/PhaGen", trust_remote_code=True)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Tim419/PhaGen", trust_remote_code=True, device_map="auto")
Quick Links

PhaGen

PhaGen is the frozen paper model: 152,421,601 parameters, three hierarchical stages with dimensions 512/256/196, maximum token canvas 131,072, and a nine-token character-level DNA tokenizer.

This repository contains the standalone inference weights, tokenizer, model configuration and custom code. The published PhaGen weights are identified by their Hub revision and SHA-256 checksum in the release manifest.

Load

Use PyTorch 2.5, Transformers 4.54.1, einops 0.8, beartype 0.22, accelerate and safetensors. The PhaGen source repository also provides an installable phagen CLI and an agent skill.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Tim419/PhaGen")
model = AutoModelForCausalLM.from_pretrained(
    "Tim419/PhaGen", trust_remote_code=True,
    attn_implementation="sdpa", torch_dtype=torch.float32,
).eval()
inputs = tokenizer("ACGTACGT", return_tensors="pt", add_special_tokens=False)
with torch.inference_mode():
    result = model(**inputs, is_causal=False)
print(result.logits.shape)

Pin a Hub commit hash with revision for repeatable loading. Custom model code must be reviewed before enabling trust_remote_code; the installed PhaGen CLI instead uses its own package implementation and verifies the frozen weight checksum.

Outputs and interpretation

Forward returns logits and three hierarchical hidden-state tensors. Diffusion generation uses the supplied MDMGenerationConfig from generation_utils.py with a masked canvas. The nominal token budget is distinct from extracted DNA lengths. Sequence scores from a masked model are not automatically exact autoregressive likelihoods.

Training and limitations

Recorded training commit: 13137c158fa33866d9bcc28c8af04f226f4c4d7e. Runtime records indicate learning rate 2e-5, global batch 72, seed 42 and mixed precision. The prior warm-start checkpoint is unavailable, and an exact runtime dirty-tree source snapshot was not retained. Exact retraining to this checkpoint is not established.

The historical train/validation split has disjoint IDs but 549 shared exact sequence contents, affecting 553 validation records (11.10%). This release preserves the paper artifact; it does not claim an independent historical validation partition.

The hierarchy follows the megaDNA research lineage; training uses the VeOmni framework, and diffusion utilities describe adaptation from Dream. See NOTICE.md for attribution and unresolved upstream license provenance. No blanket license claim for third-party source or biological datasets is made by this model card.

These are computational model outputs, not experimentally validated biological function. The selected model SHA-256 and source hashes are recorded in release_manifest.json.

Downloads last month
22
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support