aether-55m / README.md
ameyg910's picture
docs: model card
3cc40d9 verified
|
Raw
History Blame Contribute Delete
3.65 kB
---
license: apache-2.0
library_name: aether
tags:
- diffusion
- masked-diffusion
- language-model
- mdlm
datasets:
- wikitext
pipeline_tag: text-generation
---
# ameyg910/aether-55m
A masked (absorbing-state) diffusion language model trained with
[Aether](https://github.com/ameyg910/aether).
Unlike an autoregressive model, generation does not proceed left to right. The
model starts from an all-`[MASK]` sequence and unmasks progressively, so the
number of forward passes (**NFE**) is a knob rather than a function of sequence
length.
## Model details
| | |
| --- | --- |
| Architecture | bidirectional DiT denoiser, AdaLN-Zero time conditioning |
| Objective | MDLM / SUBS masked-diffusion loss |
| Parameters | 55,543,634 |
| Width / depth / heads | 384 / 6 / 6 |
| Context length | 1024 |
| Vocabulary | 50,258 (GPT-2 + `[MASK]`) |
| Training steps | 30,000 |
| Tokens seen | 3,932,160,000 |
| License | Apache-2.0 |
## Evaluation
Measured with `aether-eval`; see
[the evaluation protocol](https://github.com/ameyg910/aether/blob/main/docs/evaluation.md).
| metric | value |
| --- | --- |
| NELBO (nats/token) | 7.136 |
| Bits per dim | 10.3 |
| Perplexity (upper bound) | 1257 |
| MAUVE | 0.999 |
| distinct-2 | 0.971 |
| Sampler / steps | ancestral / 128 |
> **Perplexity here is an upper bound, not an exact likelihood.** A masked
> diffusion model has no exact factorization of `log p(x)`; what is reported is a
> Monte Carlo estimate of a variational bound. It is comparable to other diffusion
> models evaluated the same way, and **not** directly comparable to an
> autoregressive model's exact perplexity, which would flatter the AR model.
## Usage
Serve it:
```bash
pip install "aether-dlm[serve]"
aether-serve serve.model_version=hf:ameyg910/aether-55m@v1.0.0
curl -X POST localhost:8000/generate \
-H 'content-type: application/json' \
-d '{"n_samples":2,"length":64,"steps":64,"sampler":"ancestral"}'
```
Or load it directly:
```python
import torch
from huggingface_hub import hf_hub_download
from aether.models.loading import build_model_from_checkpoint
from aether.diffusion.samplers import sample
path = hf_hub_download("ameyg910/aether-55m", "latest.pt", revision="v1.0.0")
model, config = build_model_from_checkpoint(
torch.load(path, map_location="cpu", weights_only=False)
)
out = sample(model.eval(), batch=2, length=64,
mask_token_id=config.vocab_size - 1, steps=64)
print(out.tokens.shape, "NFE:", out.nfe)
```
## Intended use
Research and demonstration of masked diffusion language modelling: studying the
NFE-quality tradeoff, comparing sampling strategies, and as a fixture for
inference-serving work.
## Limitations
- **Small and undertrained.** 55,543,634 parameters and roughly
3,932,160,000 tokens. It captures vocabulary and local phrasing, not
long-range coherence or factual grounding.
- **Unconditional.** There is no prompt input; it generates from an all-`[MASK]`
sequence. Prompt-conditioned infilling is a natural extension the architecture
supports but this release does not implement.
- **No alignment of any kind.** No instruction tuning, no safety filtering, no
RLHF. Output may be offensive, false, or nonsensical.
- **Inherits its corpus.** Trained on wikitext, and reproduces the biases and
errors in it.
Not suitable for production text generation, question answering, or any use where
output correctness matters.
## Citation
```bibtex
@software{aether,
author = {Gupta, Amey},
title = {Aether: a production platform for masked diffusion language models},
year = {2026},
url = {https://github.com/ameyg910/aether}
}
```