# Kimi K3 — Architecture Notes > Deep dive on the architecture of **Kimi K3: Open Frontier Intelligence** (Kimi Team, Moonshot AI, > July 2026). Source: the 47-page technical report. Claims marked **[report]** are stated there; > **[analysis]** is our inference. The organizing idea is stated plainly in the report and it is worth taking literally: **a Transformer moves information along three axes — sequence length, network depth, and model width — and K3 replaces one component on each axis.** | Axis | Kimi K2 | Kimi K3 | What it fixes | |---|---|---|---| | Sequence (token mixing) | All MLA | **Hybrid: 3× KDA + 1× Gated MLA**, repeated | Compute/memory at 1M context | | Depth (layer mixing) | Standard residual accumulation | **Attention Residuals (Block AttnRes)** | Deep layers only see a compressed history | | Width (channel mixing) | DeepSeekMoE, 384 experts | **Stable LatentMoE, 896 experts / 16 active** | Expert specialization space vs. communication and stability | | Vision | — | **MoonViT-V2, 401M, trained from scratch** | Joint-optimization instability | | Optimizer | Muon | **Per-Head Muon** | Uneven update scale across heads | Together with refined data and training recipes these deliver an approximately **2.5× improvement in overall scaling efficiency** over K2 — same FLOPs, lower held-out validation loss. The report does **not** decompose how much of that comes from architecture vs. data vs. recipe. **[report]** --- ## 0. Specification diff, K2 → K3 | | Kimi K2 | Kimi K3 | Δ | |---|---|---|---| | Layers | 61 | **93** | ↑52% | | Total parameters | 1.04T | **2.78T** | ↑167% | | Activated parameters | 32.6B | **104.2B** | ↑220% | | Hidden dimension | 7,168 | 7,168 | = | | Latent MoE dimension | — | **3,584 (0.5×)** | new | | MoE hidden dim per expert | 2,048 | 3,072 | ↑50% | | Routed experts | 384 | **896** | ↑133% | | Experts active per token | 8 | **16** | ↑100% | | Shared experts | 1 | **2** | ↑100% | | Attention heads | 64 | 96 | ↑50% | | Dense layers | 1 | 1 | = | | Vocabulary | 160K | 160K | = | | Training context length | 128K | **1M** | 8× | | Attention mechanism | MLA | **Hybrid KDA–MLA** | — | | Activation | SwiGLU | **SiTU-GLU** | — | | Attention-layer composition | 61 MLA | **69 KDA + 24 MLA** | — | | MTP layers | 1 | 1 | = | | ViT params / layers / patch / heads | — | **401M / 27 / 14 / 12** | new | Sparsity is 896/16 = **56**. That ratio is the central tension of the design: a bigger expert pool with more experts active per token buys specialization, but communication and expert-weight traffic grow with routing multiplicity. LatentMoE exists to make that trade affordable (§3). --- ## 1. Hybrid attention Each block is **3 KDA layers followed by 1 Gated MLA layer** — a 3:1 ratio, repeated throughout the backbone. An **additional Gated MLA layer sits at the end of the backbone**, guaranteeing the final layer always performs global attention. ### 1.1 Kimi Delta Attention KDA comes from Kimi Linear ([arXiv:2510.26692](https://arxiv.org/abs/2510.26692), MIT, with released 48B-A3B checkpoints). It is a **delta-rule recurrence with a channel-wise forget gate**: ``` S_t = (I − β_t k_t k_tᵀ) Diag(α_t) S_{t−1} + β_t k_t v_tᵀ ō_t = S_tᵀ q_t ``` - `α_t ∈ (0,1)^{d_k}` — a **channel-wise** one-step retention factor. Not a scalar: every channel decides its own forgetting rate. - `β_t ∈ (0,1)` — delta-rule write strength. - q/k/v projections apply ShortConv then Swish; q and k are further L2-normalized. The state `S ∈ R^{d_k×d_v}` is **fixed size** — it replaces a KV cache that grows with sequence length. That is the physical basis for 1M context. ### 1.2 The lower-bounded decay — the most instructive decision in the report This is worth studying not for the result but for the *shape* of the reasoning: **the algorithm did not change; only the range of one quantity was tightened, and that bought a dense Tensor Core path.** **The problem.** KDA is recurrent across chunks and parallel within each chunk. The chunkwise form rescales keys by the reciprocal cumulative decay `1/Γ^{1→C}`. Since `Γ` is a product of retention factors in `(0,1)`, this reciprocal **grows without bound and overflows in finite precision**. **Kimi Linear's answer.** Control the numerical range by computing relative decay in log space and splitting each chunk into secondary 16-token tiles. Off-diagonal tiles can then use dense Tensor Core matmuls — but **diagonal tiles still require explicit position-pair computation**, which becomes the main intra-chunk bottleneck. **K3's answer.** Change the mapping from decay logits `z` to per-step log-decay `g`: | | Mapping | Range | |---|---|---| | GDN / Mamba-2 / Kimi Linear | `g = −e^{A_h} · Softplus(z)` | `(−∞, 0)` — unbounded | | **Kimi K3** | `g = g_min · Sigmoid(e^{A_h} z)` | `(g_min, 0)` — **bounded, `g_min = −5` fixed** | Then `α = exp(g) ∈ (e^{g_min}, 1)`. With `g_min = −5`, every retention factor satisfies `α ≥ e^{−5} ≈ 6.7×10⁻³`, so the cumulative log-decay over a **16-token tile lies in `(−80, 0)`**, the reciprocal rescaling factor is **smaller than `e^80`**, and that **remains within the BF16 dynamic range**. **Consequence:** both diagonal and off-diagonal causal tiles now use dense Tensor Core matmuls. The position-pair diagonal path is **eliminated entirely**. `A_h` is a learnable per-head log-scale initialized to 0; `b_α` follows Kimi Linear's initialization. The report notes this parameterization is closely related to lower-bounded recurrence gates in prior work (HGRN2, Griffin, RWKV-7). > **[analysis]** The transferable lesson is diagnostic: whenever a kernel needs a special-cased > branch, an fp32 fallback, or a separate code path "for numerical reasons", the first question is > whether the algorithm genuinely requires it — or whether some freely-chosen parameterization simply > left a range wider than it needed to be. ### 1.3 Full-rank output gate Kimi Linear used a low-rank output gate. K3 replaces it with an **input-dependent full-rank projection**, applied after head-wise RMSNorm of the recurrent output: ``` y = W_o [ Sigmoid(W_g x) ⊙ RMSNorm(ō) ] ``` ### 1.4 Gated MLA, and NoPE MLA (from DeepSeek-V2) compresses each token's KV into a low-dimensional latent `c = W_c x`, caching `c` and reconstructing keys and values through learned up-projections during attention. K3 retains it in the periodic global-attention layers and adds the same **input-dependent, channel-wise full-rank output gate**. `W_g` is full rank, matching KDA's new parameterization; the gate lets each token modulate which channels it reads from global attention. **All MLA layers use NoPE — no positional encoding at all.** Positional information is carried implicitly by KDA's recurrence and decay. Two consequences: 1. **Context extension requires no positional-encoding changes** — no RoPE base retuning, no YaRN interpolation. The model extrapolates directly to 1M tokens. 2. It cleanly divides labor: KDA layers provide position-sensitive, recency-aware mixing; MLA layers provide unrestricted, position-agnostic global content interaction. **One low-precision detail.** To correct the biased rounding error that arises in flash attention, K3 adopts the method of [arXiv:2510.04212](https://arxiv.org/abs/2510.04212) and **keeps the attention output in FP32 during training**. That doubles the on-chip footprint of the output tile — so the training kernel was redesigned to **overlap it with the KV staging buffers instead of the query tile**, freeing shared memory for a deeper KV pipeline and higher throughput. --- ## 2. Attention Residuals — attention applied to the depth axis **The motivation, stated as an analogy worth remembering:** standard residual connections compress all prior information into a single state `h_l` over depth — *a bottleneck reminiscent of RNNs over time*. Transformers replaced recurrence over time with attention, letting each position selectively access all previous positions with data-dependent weights. AttnRes applies the same methodology to depth: **each layer selectively retrieves representations from all preceding layers** rather than accumulating them uniformly. ### 2.1 Full Attention Residuals Each layer `l` gets a learnable pseudo-query `q_l = w_l ∈ R^d`; keys and values are the outputs of all preceding layers, with `i = 0` being the token embedding: ``` φ(q, k) = exp( qᵀ RMSNorm(k) ) ← RMSNorm prevents large-magnitude layers dominating α_{i→l} = φ(q_l, k_i) / Σ_j φ(q_l, k_j) h_l = Σ_{i=0}^{l−1} α_{i→l} · v_i ``` Since depth is modest (`L < 100`), the `O(L²d)` arithmetic is affordable. **The practical overhead is the `O(Ld)` memory** of keeping all layer outputs alive — plus cross-stage communication under pipeline parallelism. ### 2.2 Block Attention Residuals — what K3 actually uses Partition `L` layers into `N` blocks of `S = L/N` layers. - **Within a block**, layer outputs are reduced to a single block representation by summation: `b_n = Σ_{j∈B_n} f_j(h_j)`, with `b_0 = h_1` so the token embedding is always a source. - **Across blocks**, full attention is applied over only the `N` block-level representations. Memory and communication overhead drop from `O(Ld)` to `O(Nd)`. **K3's configuration: 8 blocks of 12 layers** — giving a partial final block, and 9 total blocks when counting the embedding layer. The report cites empirical evidence that `N ≈ 8` recovers most of the benefit across model scales. The block structure also **bounds inference-time state**, and lets the parallel inter-block results be merged with the sequential intra-block partial sums via **online softmax**, significantly reducing inference-time cost. --- ## 3. Stable LatentMoE — three patches that make 896 experts work **LatentMoE** ([arXiv:2601.18089](https://arxiv.org/abs/2601.18089)) separates the model width seen by the routed experts from the full width: **shared experts retain a full-width path for common transformations, while specialized routed experts operate in a compact latent space of width `ℓ`.** ``` u = Σ_{i∈T_k(x)} p_i · E_i^routed(W↓ x) ← routed path at ℓ = 3584 (0.5 × 7168) y = Σ_{j=1}^{N_s} E_j^shared(x) + W↑ RMSNorm(u) ← shared path at full width d ``` K3 fixes `N_s = 2` full-width shared experts per layer. This is what makes scaling channel mixing to 896 routed experts with 16 active per token affordable in communication and weight traffic. But extreme sparsity amplifies two failure modes of the vanilla design, and each patch targets one. ### 3.1 Patch 1 — RMSNorm before the up-projection (Normalized LatentMoE) **Failure mode:** the routed path composes `W↓`, a gated multi-branch expert FFN, and `W↑` into a chain of nearly four consecutive matrix multiplications. That ill-conditioned structure at 2.78T scale **produces exploding internal activations in the routed branch**. Original LatentMoE applies `W↑` directly to the aggregated routed representation `u`, whose scale varies with the selected experts and their routing weights. **Patch:** insert an RMSNorm **between expert aggregation and the up-projection**. This reduces the sensitivity of the routed branch to scale variation before it is combined with the full-width shared branch. Beyond stabilizing training, the report states the additional RMSNorm **consistently improves validation loss and downstream benchmarks**. ### 3.2 Patch 2 — SiTU-GLU (Sigmoid Tanh Unit GLU) **Failure mode:** in SwiGLU **both multiplicative factors are unbounded**, so coincident large coordinates produce activation outliers and raise overflow risk in low-precision arithmetic. The original GLU's sigmoid gate avoids unbounded gate growth, but discards the approximately-linear positive regime of Swish that makes it work. **Patch:** apply the smooth cap `softcap(x, β) = β·tanh(x/β)` to the linear factor of the Swish gate and, independently, to the up branch: ``` SiTU-GLU(x) = [ β₁ tanh(W_g x / β₁) ⊙ Sigmoid(W_g x) ] ⊙ [ β₂ tanh(W_u x / β₂) ] ``` | Property | Value | |---|---| | Hyperparameters | **β₁ = 4** (gate branch), **β₂ = 25** (up branch) | | Output bound | `‖SiTU-GLU(x)‖_∞ ≤ β₁β₂ = 100` | | Near the origin | `β tanh(z/β) = z + O(z³/β²)` — **matches SwiGLU to first order** | | Limit | Recovers SwiGLU pointwise as `β₁, β₂ → ∞` | Unlike hard clamping of gate pre-activations, **the smooth cap preserves nonzero gradients away from saturation boundaries**, which the report finds gives better training behavior. ### 3.3 Patch 3 — Quantile Balancing (QB) K3 uses **auxiliary-loss-free routing**: an expert-specific bias `b_j` is added to the router score used for Top-k selection, but omitted from the mixture weights. ``` T_i = argtop_k(s_i + b), p_{i,j} = s_{i,j} / Σ_{r∈T_i} s_{i,r}, j ∈ T_i ``` Because `b` is omitted from `p`, it **regulates dispatch without altering mixture weights or the gradient-based optimization of the router**. **Failure mode:** the original method updates `b` with a fixed-step sign rule `b_j ← b_j + γ·sign(ℓ̄ − ℓ_j)`. Maintaining balanced loads **becomes much harder as LatentMoE grows the pool to 896 experts per layer**: `γ` trades off slow adaptation against oscillation, imbalanced routing slows expert-parallel training, and some experts may end up poorly trained. **Patch — set each bias from the router-score quantile that matches its target load.** With target load `q = mk/n` for a batch of `m` tokens over `n` experts: 1. Replace Top-k selection with **Top-(k+1)** on the biased score. The first `k` entries are the routes actually taken; the `(k+1)`-th entry is the cutoff `α_i` that an expert must exceed to enter token `i`'s Top-k. Taking the cutoff from Top-(k+1) routing **avoids a separate token-side quantile pass**. 2. With cutoffs fixed, the count of tokens routed to expert `j` under candidate bias `b̂_j` is monotonically decreasing in the threshold `−b̂_j`. Setting that count to `q` makes `−b̂_j` the `(q+1)`-th largest margin `s_{i,j} − α_i`. Since `q/m = k/n`, this is the `(1 − k/n)`-quantile: ``` b̂_j ← − quantile_{1−k/n}( s_{:,j} − α ) b ← b̂ − mean(b̂) · 1 ← mean-centering ``` The update **takes effect only at the next step** — a batch is never routed with a bias derived from itself. The final bias is **frozen at inference**. **Why it is principled.** Appendix C derives QB from the maximum-score balanced assignment problem. The LP relaxation is exact (total unimodularity of the bipartite b-matching polytope); the convex dual is minimized by **alternating exact coordinate minimization**, and both subproblems turn out to be quantiles along the token and expert axes respectively — hence the name. The original sign-based loss-free update is recovered as a **SignSGD step on that same dual objective**; QB jumps directly to the exact coordinate minimizer. In the reported experiments it **equilibrates within a few update steps even for nearly 10³ experts**. **Making the quantile computable at scale (Appendix D).** The quantile spans the whole global batch — millions of margins sharded across ranks and gradient-accumulation steps — so gathering them exactly is not viable inside the training loop. K3 instead maintains a **binned histogram per expert**: - Each rank scatter-adds its local values into a per-expert count matrix `H ∈ N^{n×B}` during the forward pass, accumulating over micro-batches with **no communication**. A single all-reduce at the end of the step sums local counts into the global histogram. - **`B = 1000` bins** gives an error of **at most a few 10⁻³**, with no measurable residual load imbalance observed. - Communication is one integer all-reduce of `n × B` values per layer per step — **independent of `m`**, and in their configuration **below 1% of the cost** of exchanging raw margins. - Because counts are additive, the estimate is **exactly invariant to how tokens are partitioned across ranks**: it is the quantile of the *pooled* global batch, not an average of per-rank quantiles — which generally differ. --- ## 4. MoonViT-V2 — the vision tower trained from scratch **The departure:** prior practice, including Kimi K2.5 itself, initializes the vision encoder from a contrastively pre-trained model such as SigLIP, on the premise that pre-trained visual knowledge gives the model a head start. **K3 trains MoonViT-V2 entirely from scratch with next-token prediction.** Three reasons and results, per the report: 1. **Training stability (the primary motivation).** When a pre-trained encoder is attached to the LLM, joint optimization becomes unstable. Report Fig. 6 is the evidence: the SigLIP-initialized MoonViT-3D shows **persistently higher vision-tower gradient norms with frequent spikes**, while MoonViT-V2 remains stable throughout training. 2. **Objective alignment.** Next-token prediction lets the encoder's representations be **shaped directly by the language-modeling objective**, rather than by a contrastive loss that favors global semantics over fine-grained textual and structural cues. 3. **Result.** MoonViT-V2 **matches the SigLIP-initialized baseline across vision evaluations** — indicating, in the report's words, that contrastive pre-training is unnecessary as an initialization for multimodal language models at scale. > **[analysis]** This is evidence *at scale*, with a full multimodal corpus and a 2.8T backbone. It is > not an argument that from-scratch vision towers win in small-scale or fine-tuning regimes, where the > head start from pre-trained visual knowledge is likely still real. What generalizes more safely is > the **diagnostic**: if you observe persistent vision-tower gradient spikes during joint optimization, > this report says the root cause may be an objective mismatch between contrastive initialization and > the language-modeling objective — not your learning rate. That makes vision-tower gradient norm a > monitoring signal worth having. **Architecture.** - 27-layer ViT, **401M parameters**, adopting **RMSNorm and removing all bias terms** from its linear and attention projections — a design that further stabilizes from-scratch optimization. - **Images and videos are processed with fully shared parameters.** Attention is factorized into intra-frame spatial and inter-frame temporal passes; temporal pooling further compresses tokens along the time dimension. - Before projection, a **2×2 pixel-shuffle downsampling** reduces visual token count by 4×, keeping inputs of up to **3584 × 3584 pixels** affordable within the 1M-token context. - A lightweight **MLP projector** maps visual features into the shared embedding space. **Why native multimodality matters architecturally.** Text, images and video are processed by a single shared backbone within one context, with no post-hoc modality-alignment stage. **Rendered outputs and the code that produced them live in the same token stream** — the model can write code, inspect screenshots, and iteratively refine visual artifacts (UIs, graphics, video) **with no cross-model hand-off**. This is the architectural precondition for the video-editing and motion-design results in report §7. --- ## 5. Per-Head Muon K3 follows K2 in using **Muon** for matrix parameters, refined into a **per-head variant for attention projections**: instead of applying Newton–Schulz orthogonalization to the full Q, K, V projection matrices, their momentum matrices are **partitioned along the head dimension and each head's block is orthogonalized separately**. **The intuition:** full-matrix orthogonalization treats all heads as a single coupled block, so heads with larger gradient or momentum scales dominate the shared update direction while smaller-scale heads receive insufficiently normalized updates. **Per-head orthogonalization equalizes the update scale across heads.** In practice this yields more balanced learning dynamics across heads and improves training stability at larger scales. It also **slightly reduces optimizer overhead**, since Newton–Schulz iterations on tall per-head blocks are cheaper than on the full projection matrix. --- ## References - Kimi K3 technical report · [blog](https://www.kimi.com/blog/kimi-k3) · [weights](https://huggingface.co/moonshotai/Kimi-K3) - Kimi Linear (KDA) [arXiv:2510.26692](https://arxiv.org/abs/2510.26692) · [code](https://github.com/MoonshotAI/Kimi-Linear) - LatentMoE [arXiv:2601.18089](https://arxiv.org/abs/2601.18089) · DeepSeekMoE [arXiv:2401.06066](https://arxiv.org/abs/2401.06066) · DeepSeek-V2 / MLA [arXiv:2405.04434](https://arxiv.org/abs/2405.04434) - Gated DeltaNet [ICLR 2025](https://openreview.net/forum?id=r8H7xhYPwz) · Mamba-2 / SSD [arXiv:2405.21060](https://arxiv.org/abs/2405.21060) - Gated Attention [arXiv:2505.06708](https://arxiv.org/abs/2505.06708) · GLU Variants [arXiv:2002.05202](https://arxiv.org/abs/2002.05202) · PowLU [arXiv:2605.25704](https://arxiv.org/abs/2605.25704) - Low-precision flash attention rounding [arXiv:2510.04212](https://arxiv.org/abs/2510.04212) - Muon [kellerjordan.github.io/posts/muon](https://kellerjordan.github.io/posts/muon/) · Muon is Scalable [arXiv:2502.16982](https://arxiv.org/abs/2502.16982) - Load balancing: auxiliary-loss-free (DeepSeek-V3) [arXiv:2412.19437](https://arxiv.org/abs/2412.19437) · BIP [arXiv:2502.15451](https://arxiv.org/abs/2502.15451) · the quantile view, in Chinese [spaces.ac.cn/archives/11619](https://spaces.ac.cn/archives/11619)