--- license: mit language: - en - de - ru - ja tags: - pytorch - text-generation - mixture-of-experts - linear-attention - hybrid-attention - from-scratch - research datasets: - HuggingFaceFW/fineweb-edu - HuggingFaceTB/finemath - HuggingFaceFW/fineweb-2 - bigcode/starcoderdata pipeline_tag: text-generation --- # Meridian-Mini > **Status: training in progress.** Weights will be uploaded when the run finishes. This card describes the plan and will be updated with results. A 236M-parameter mixture-of-experts language model with hybrid linear/full attention, trained from scratch on a single RTX 3060 (12GB). Meridian-Mini is the full-size sibling of [Meridian-Tiny](https://huggingface.co/YOUR_USERNAME/Meridian-Tiny), the 12M-parameter smoke test that validated the code, data pipeline and training loop. Same architecture, same code, same data; roughly 20× more parameters, 3× more depth and 50× more training tokens. The goal is to find out whether a stack of recent architecture ideas can be combined, scaled down and trained on consumer hardware, and whether the components that showed no measurable effect at 4 layers start to matter at 12. ## Architecture Each layer is built from up to three sublayers, each wrapped in its own hyper-connection: 1. **Engram** (layers 2 and 7): hashed n-gram memory 2. **Mixer**: KDA linear attention in 9 layers, MLA full attention in layers 4, 8 and 12 3. **FFN**: dense SwiGLU in layer 1, latent MoE in the other 11 | Component | What it does | Origin | |---|---|---| | **KDA** (Kimi Delta Attention) | Linear attention with a fixed-size memory matrix per head, updated by a gated delta rule with per-channel forgetting. Includes a causal short convolution. | Kimi Linear (Moonshot AI) | | **MLA** (Multi-head Latent Attention) | Full attention with queries and keys/values compressed through low-rank latents, plus a decoupled RoPE key shared across heads and a sigmoid output gate. | DeepSeek-V2/V3; output gate as in Qwen3-Next | | **Latent MoE** | Tokens are projected into a smaller latent space before the routed experts. Sigmoid router, grouped top-k routing, one always-on shared expert, and aux-loss-free load balancing via a per-expert routing bias. | DeepSeek-V3 routing; latent-space experts | | **mHC** (manifold-constrained hyper-connections) | The residual stream is widened to 4 parallel streams. Each sublayer learns dynamic pre/post weights and a stream-mixing matrix constrained to be doubly stochastic via Sinkhorn iterations. | DeepSeek | | **Block AttnRes** (attention residuals) | At block boundaries (every 4 layers), the residual state is rebuilt as a softmax-weighted mix of all previous block snapshots, scored by a learned pseudo-query per boundary. | Kimi (Moonshot AI) | | **Engram** | For each token, the preceding 2- and 3-grams are hashed (4 hash functions each) into a large embedding table. Retrieved memory is gated against the current hidden state before being added. | DeepSeek | See the [Meridian-Tiny card](https://huggingface.co/hashtagg1/Meridian-Tiny) for implementation notes and how this interpretation differs from the original papers. ## Model details | | Meridian-Mini | |---|---| | Parameters (total) | 236,299,134 | | Active per token | ≈78M, of which ≈45M excluding embeddings | | Hidden size | 512 | | Layers | 12 | | Attention heads | 8 | | KDA head dim | 64 | | MLA dims | q rank 256, kv rank 128, nope 64, rope 32, v 64 | | Routed experts | 32, top-4, in 4 groups (top-2 groups per token) | | Shared experts | 1 | | Expert latent / intermediate | 384 / 256 | | Dense FFN (layer 1) | 2,048 | | Hyper-connection streams | 4 | | AttnRes block size | 4 layers | | Engram | 2 layers, 524,287 rows × 64, 2- and 3-grams, 4 hash heads per order | | Context length | 2,048 | | Vocabulary | 32,768 (same tokenizer as Meridian-Tiny) | ### Where the parameters are | Part | Parameters | Share | |---|---|---| | MoE layers (11) | 112.6M | 48% | | Engram (2 layers) | 68.2M | 29% | | Token embedding + LM head (untied) | 33.6M | 14% | | Attention: KDA (9) + MLA (3) | 17.4M | 7% | | Dense FFN (layer 1) | 3.1M | 1% | | Hyper-connections, AttnRes, norms | 1.3M | <1% | Over three quarters of the parameters are **sparse**: routed experts that only fire for some tokens, and Engram table rows that are only read when their n-gram appears. Each token touches 4 of 32 experts and 8 table rows per Engram layer, so the compute per token is closer to that of a ~78M dense model. ## Training plan | | | |---|---| | Tokens | 1B (the packed corpus has 4B available for continued training) | | Steps | 7,629 | | Micro-batch | 1 × 2,048 tokens | | Gradient accumulation | 64 | | Tokens per optimizer step | 131,072 | | Optimizer | AdamW, β = (0.9, 0.95), weight decay 0.1 | | Weight decay exclusions | 1D params, token embeddings, Engram tables, hyper-connection params | | Peak LR | 6e-4 | | Schedule | WSD: 2% linear warmup, stable, 15% linear decay to zero | | Gradient clipping | 1.0 | | Precision | bf16 autocast, fp32 master weights | | MoE balancing | routing-bias update of ±1e-3 per step, no auxiliary loss | | Validation | every 250 steps, reported separately per data source | | Checkpoints | every 100 steps (~35 minutes) | | Hardware | 1× RTX 3060 12GB | | Expected wall time | ~44 hours | The WSD schedule was chosen so the run can be extended: training can continue from a stable-phase checkpoint on more of the 4B-token corpus, with the decay phase applied at the new end point. ### Training data Same corpus and mix as Meridian-Tiny: | Share | Domain | Source | |---|---|---| | 60% | English web | [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) (`sample-10BT`) | | 15% | Math | [FineMath](https://huggingface.co/datasets/HuggingFaceTB/finemath) (`finemath-4plus`) | | 7.5% | Python | [StarCoderData](https://huggingface.co/datasets/bigcode/starcoderdata) | | 4.5% | C++ | StarCoderData | | 3% | Rust | StarCoderData | | 4% | German | [FineWeb-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2) (`deu_Latn`) | | 4% | Russian | FineWeb-2 (`rus_Cyrl`) | | 2% | Japanese | FineWeb-2 (`jpn_Jpan`) | FineWeb-Edu, FineMath and FineWeb-2 are released under ODC-By 1.0. StarCoderData is derived from The Stack; code in it remains under its original licenses and is subject to the dataset's terms of use. ## Making it trainable on a 3060 The first working version of this model trained at **69 tokens per second**, which would have taken about 650 days to get through the corpus. Each fix below was profiled first, then verified against a slower reference implementation before being swapped in. | Change | Throughput (tokens/s, fwd + bwd, 1 × 2,048) | |---|---| | Naive PyTorch everything | ~70 | | KDA: per-token Python loop → `chunk_kda` kernel from [flash-linear-attention](https://github.com/fla-org/flash-linear-attention) | 2,295 | | MoE: per-expert loop → fixed-capacity buffers and batched matmuls, no CPU–GPU syncs | 4,867 | | Sinkhorn: 20 unrolled iterations → fused with `torch.compile` | **6,382** | That's roughly a **90× speedup**, with every fast path checked against its reference: - `chunk_kda` vs the hand-written recurrent KDA: max abs difference ~2e-4 - Batched MoE vs the per-expert loop (no capacity overflow): exact match - Compiled vs eager Sinkhorn: match within 1e-5 The remaining bottleneck at this size is kernel-launch overhead, not arithmetic: most of each step is spent issuing thousands of small GPU operations. Memory at batch 1 × 2,048 is about 6 GB for weights, gradients and activations, plus about 1.9 GB of AdamW state. Batch 2 fits in isolation (~11 GB) but leaves no room for the optimizer, so the larger effective batch comes from gradient accumulation. ## Open questions this run should answer **Do mHC and AttnRes matter with more depth?** At 4 layers, removing either one changed validation loss by less than run-to-run noise. Both mechanisms route information across depth, and a 12-layer model gives them three times as much to work with. mHC also carries a real speed cost, so if it still shows no benefit here, dropping it would buy substantially faster training. **Does Engram's advantage grow?** In Meridian-Tiny, Engram was the only component with a clear effect (about 4× seed noise), concentrated in code. With 50× more data, the n-gram tables for German, Russian and Japanese should become useful too. **How do the non-English languages develop?** Meridian-Tiny's German drifted into English mid-sentence, while Russian and Japanese stayed in their own scripts. More capacity and data should show whether that drift is a scale problem or a data-share problem. **Does routing specialize by domain?** With 32 experts and a mixed corpus of prose, math, three programming languages and three human languages, it should be possible to see whether experts split by domain. ## Results *To be added when training finishes: per-source validation curves, final losses, a comparison with Meridian-Tiny, and generation samples.* ## Usage *Weights not yet uploaded.* Once they are, loading will work the same way as Meridian-Tiny: download `model.py`, `presets.py`, `tokenizer.json`, `config.json` and `model.safetensors`, and load with `MiniMeridian(FULL["model"])`. Training at this size requires [`flash-linear-attention`](https://github.com/fla-org/flash-linear-attention) for the KDA kernel (plus `triton-windows` on Windows). Inference works without it, using the pure-PyTorch fallback. ## Limitations - This is a research model, not an assistant. At 1B training tokens, expect fluent-looking text with limited factual reliability and weak reasoning. - Not instruction-tuned, not safety-tuned. It continues text; it does not follow instructions. - Trained on unfiltered-beyond-source web text, and may reproduce anything found in it. - Several components are this project's own interpretation of published ideas, and the mHC + AttnRes combination is untested elsewhere. ## Acknowledgements Architecture ideas from Moonshot AI (Kimi Linear, KDA, attention residuals), DeepSeek (MLA, MoE routing, mHC, Engram) and Qwen (gated attention). Fast KDA kernel from the flash-linear-attention project. Data from Hugging Face (FineWeb-Edu, FineMath, FineWeb-2) and BigCode (StarCoderData).