ameyg910/aether-55m
A masked (absorbing-state) diffusion language model trained with Aether.
Unlike an autoregressive model, generation does not proceed left to right. The
model starts from an all-[MASK] sequence and unmasks progressively, so the
number of forward passes (NFE) is a knob rather than a function of sequence
length.
Model details
| Architecture | bidirectional DiT denoiser, AdaLN-Zero time conditioning |
| Objective | MDLM / SUBS masked-diffusion loss |
| Parameters | 55,543,634 |
| Width / depth / heads | 384 / 6 / 6 |
| Context length | 1024 |
| Vocabulary | 50,258 (GPT-2 + [MASK]) |
| Training steps | 30,000 |
| Tokens seen | 3,932,160,000 |
| License | Apache-2.0 |
Evaluation
Measured with aether-eval; see
the evaluation protocol.
| metric | value |
|---|---|
| NELBO (nats/token) | 7.136 |
| Bits per dim | 10.3 |
| Perplexity (upper bound) | 1257 |
| MAUVE | 0.999 |
| distinct-2 | 0.971 |
| Sampler / steps | ancestral / 128 |
Perplexity here is an upper bound, not an exact likelihood. A masked diffusion model has no exact factorization of
log p(x); what is reported is a Monte Carlo estimate of a variational bound. It is comparable to other diffusion models evaluated the same way, and not directly comparable to an autoregressive model's exact perplexity, which would flatter the AR model.
Usage
Serve it:
pip install "aether-dlm[serve]"
aether-serve serve.model_version=hf:ameyg910/aether-55m@v1.0.0
curl -X POST localhost:8000/generate \
-H 'content-type: application/json' \
-d '{"n_samples":2,"length":64,"steps":64,"sampler":"ancestral"}'
Or load it directly:
import torch
from huggingface_hub import hf_hub_download
from aether.models.loading import build_model_from_checkpoint
from aether.diffusion.samplers import sample
path = hf_hub_download("ameyg910/aether-55m", "latest.pt", revision="v1.0.0")
model, config = build_model_from_checkpoint(
torch.load(path, map_location="cpu", weights_only=False)
)
out = sample(model.eval(), batch=2, length=64,
mask_token_id=config.vocab_size - 1, steps=64)
print(out.tokens.shape, "NFE:", out.nfe)
Intended use
Research and demonstration of masked diffusion language modelling: studying the NFE-quality tradeoff, comparing sampling strategies, and as a fixture for inference-serving work.
Limitations
- Small and undertrained. 55,543,634 parameters and roughly 3,932,160,000 tokens. It captures vocabulary and local phrasing, not long-range coherence or factual grounding.
- Unconditional. There is no prompt input; it generates from an all-
[MASK]sequence. Prompt-conditioned infilling is a natural extension the architecture supports but this release does not implement. - No alignment of any kind. No instruction tuning, no safety filtering, no RLHF. Output may be offensive, false, or nonsensical.
- Inherits its corpus. Trained on wikitext, and reproduces the biases and errors in it.
Not suitable for production text generation, question answering, or any use where output correctness matters.
Citation
@software{aether,
author = {Gupta, Amey},
title = {Aether: a production platform for masked diffusion language models},
year = {2026},
url = {https://github.com/ameyg910/aether}
}