Instructions to use guan-wang/ESM-DCLM-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use guan-wang/ESM-DCLM-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="guan-wang/ESM-DCLM-1B", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("guan-wang/ESM-DCLM-1B", trust_remote_code=True) model = AutoModelForMaskedLM.from_pretrained("guan-wang/ESM-DCLM-1B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In UNKNOWN_FILENAME: "auto_map.AutoTokenizer" must be a string
ESM-DCLM-1B
OpenESM 1B model (variant d26) trained on
DCLM. This repository contains a Hugging Face Transformers-compatible
export for the pretraining checkpoint. It uses OpenESM's custom architecture
and remote model code; it is not the built-in transformers.EsmModel
architecture.
How to Use
Install the runtime dependencies:
pip install torch transformers
Load the tokenizer and model with the custom OpenESM code enabled:
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
repo_id = "guan-wang/ESM-DCLM-1B"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained(
repo_id,
trust_remote_code=True,
)
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
inputs = tokenizer("Hello world", return_tensors="pt").to(device)
with torch.inference_mode():
outputs = model(**inputs)
print(outputs.logits.shape)
The first load may prompt you to review and allow the repository's custom model
code. Loading remote Python code requires trust_remote_code=True.
Model Details
| Detail | Value |
|---|---|
| Model family | OpenESM energy-based language model |
| Architecture | Custom OpenESM implementation (d26) |
| Parameter scale | 1B |
| Training dataset | DCLM |
| Training stage | pretraining |
| Transformer blocks | 26 |
| Embedding dimension | 1664 |
| Attention heads | 13 |
| Context length | 2048 |
| Vocabulary size | 32768 |
The parameter scale is the label used for this model in the OpenESM model configuration. The architecture and sequence settings above are read from the exported checkpoint configuration.
Files
config.json: model configuration.model*.safetensors: model weights in the standard Transformers format.modeling_esm.py: remote model and tokenizer implementation.configuration_esm.py: remote configuration implementation.tokenizer.pkl: serialized ESM tokenizer.tokenizer_config.json: tokenizer auto-loading configuration.token_bytes.pt: token byte table used by OpenESM metrics.
The original Lightning .ckpt is converted before upload and is not needed to
load this repository. Training token counts are intentionally omitted from the
repository name and file names.
Checkpoint Metadata
ebm-dclm-d26-10b
EBM pretrained on DCLM; d26, 10B.
This repository contains the following PyTorch checkpoint:
- Archive filename:
final-s=step=9499-d26-ctx2048-lr0.0012-bs1x32-muon_adamw-valid_loss=valid_loss=2.4950.ckpt - Original filename:
s=step=9499-d26-ctx2048-lr0.0012-bs1x32-muon_adamw-valid_loss=valid_loss=2.4950.ckpt - Original relative path:
pretrain/ebm/dclm/d26/10B/s=step=9499-d26-ctx2048-lr0.0012-bs1x32-muon_adamw-valid_loss=valid_loss=2.4950.ckpt
The checkpoint is uploaded as-is from the local training output. See the filename and training project for the exact architecture and loading code.
License
Please add the applicable model/data license before publishing this repo.
The code is maintained in the OpenESM repository.
- Downloads last month
- -