lfm2-quantum-128m / README.md
MarkChenX's picture
Add model architecture (nanochat/), config.json, and inference.py
68a7bb4 verified
|
Raw
History Blame Contribute Delete
2.54 kB
---
license: mit
---
# lfm2-quantum-128m (base checkpoint, step 2162)
LFM2-style hybrid quantum GPT, budget tier, trained with
[Quantum-GPT](https://github.com) `runs/run_lfm2_quantum_mini.sh`.
This is a **base** (pretrained, not instruction-tuned) checkpoint at step 2162
(final step of pretraining) -- it does raw text continuation, not chat.
## Architecture
- 16 layers, hybrid mixer: 10 gated short-conv blocks + 6 GQA attention blocks
(conv-first, pattern `CCCCCCCCCCAAAAAA`)
- `n_embd=1024`, 16 query heads / 8 KV heads (GQA), `head_dim=64`, `seq_len=1024`
- `vocab_size=32768`
- Quantum feed-forward network in every block (4 qubits, circuit depth 2,
simulated exactly in PyTorch -- see `nanochat/gpt.py`'s `QuantumMLP`)
- RoPE base `theta=1e6`, full (non-windowed) attention (`window_pattern=L`)
- Value embeddings disabled
Full config: `config.json` (metadata only) and `meta_002162.json` (exact
training config this checkpoint was produced with).
This is a **custom architecture**, not a `transformers` model -- there is no
`AutoModel` support. `config.json` is provided for discoverability/metadata
only; to actually load the model, use the bundled `nanochat/` package as
shown below.
## Inference
This repo bundles the minimal `nanochat/` source needed to load and run the
model, so it's self-contained (no need to clone the full training repo).
```bash
pip install torch tiktoken rustbpe filelock kernels
# download this repo, e.g.:
hf download MarkChenX/lfm2-quantum-128m --local-dir ./lfm2-quantum-128m
cd lfm2-quantum-128m
python inference.py --prompt "The history of quantum computing"
```
See `inference.py` for the full loading + generation code (~15 lines): it
calls `nanochat.checkpoint_manager.build_model(".", step=2162, ...)` to build
the model from `model_002162.pt` + `meta_002162.json`, loads the tokenizer
from `tokenizer/`, and streams tokens via the model's own `.generate()`.
## Resuming pretraining
`optim_002162_rank0.pt` is the matching Muon/AdamW optimizer state (momentum
buffers etc.) for this step. To continue pretraining with the original
[Quantum-GPT](https://github.com) repo, place `model_002162.pt`,
`optim_002162_rank0.pt` and `meta_002162.json` under
`$NANOCHAT_BASE_DIR/base_checkpoints/lfm2-quantum-128m/`, then run:
```bash
RESUME_FROM_STEP=2162 MODEL_TAG=lfm2-quantum-128m bash runs/run_lfm2_quantum_mini.sh
```
For SFT/RL instead of continued pretraining, only `model_002162.pt` +
`meta_002162.json` + `tokenizer/` are needed (the optimizer shard is
pretraining-only).