| # CMF v2 — Format Specification |
|
|
| *Languages: **English** · [Русский](SPEC.ru.md) · [中文](SPEC.zh.md)* |
|
|
| **Cortiq Model Format** — a single file carrying everything needed for |
| sparse, task-routed inference: quantized weights, tokenizer, per-task |
| masks, a precomputed sparse index — and, uniquely, a **swarm of skills** |
| sharing one backbone (Patent 15). |
|
|
| > Normative source: this document. Reference |
| > implementations: Rust reader/runtime (`crates/cortiq-core`, |
| > `crates/cortiq-engine`), Python writer (`converter/`), and a |
| > standalone Python reader (`python/cmf_reader.py`, stdlib + numpy). |
| |
| Three requirements, in priority order: |
| |
| 1. **Correct.** No silent corruption modes: strict magic, version, |
| `required_features`, bounds on every section, a 64-bit hash for every |
| tensor. A file is either valid or open() returns an error — there is |
| no third state. |
| 2. **Fast.** The weight section is page-aligned for mmap, every tensor |
| is 64-byte aligned (zero-copy SIMD), the tensor directory is binary — |
| read without parsing. Cold (masked-out) weights cost no RSS. |
| 3. **Compact.** Masks are bit-packed (1 bit per neuron), weights are |
| q4/q8/variable-bit, the whole file is addressed by one 128-byte |
| envelope. |
|
|
| Harmony comes not from feature count but from **a single canon**: one |
| layout per level (envelope, directory, quant block, mask), byte-for-byte |
| compatible with the validated `.vmfc` v2 format where the domains |
| overlap (tensor directory, quant layouts, `hash64`). Never two |
| definitions of the same thing. |
|
|
| Physical basis (VMF): the model is a vacuum condensate 𝒲; a skill is its |
| regular core above a critical density; a task mask selects an active |
| subset without changing weights. The format carries the consequences of |
| that physics (two-field 𝒲×θ quantization, Born importance, critical mask |
| threshold) — but **only those confirmed by measurement**. |
|
|
| --- |
|
|
| ## 1. Envelope (fixed 128 bytes) |
|
|
| All integers are little-endian. |
|
|
| ``` |
| [0x00 : 0x04] magic = b"CMF\x01" (4 bytes) |
| [0x04 : 0x08] version : u32 = 2 |
| [0x08 : 0x0C] flags : u32 (reserved, 0) |
| [0x0C : 0x10] required_features : u32 (bitmask, §1.1) |
| [0x10 : 0x18] header_off : u64 (= 128) |
| [0x18 : 0x20] header_len : u64 — JSON header (§2) |
| [0x20 : 0x28] dir_off : u64 — tensor directory (§3) |
| [0x28 : 0x30] dir_len : u64 |
| [0x30 : 0x38] data_off : u64 — weight blob; multiple of 4096 (§4) |
| [0x38 : 0x40] data_len : u64 |
| [0x40 : 0x48] masks_off : u64 — masks section (§5); 0 = absent |
| [0x48 : 0x50] masks_len : u64 |
| [0x50 : 0x58] vocab_off : u64 — tokenizer (§6); 0 = absent |
| [0x58 : 0x60] vocab_len : u64 |
| [0x60 : 0x68] index_off : u64 — sparse index (§7); 0 = absent |
| [0x68 : 0x70] index_len : u64 |
| [0x70 : 0x80] reserved : 16 bytes (§8.1: header/dir hashes) |
| ``` |
|
|
| Section order on disk: envelope → header JSON → directory → **weight |
| blob (aligned to 4096)** → masks → vocab → sparse index. A reader MUST |
| address sections ONLY through the envelope, never by assuming order. |
|
|
| ### 1.1 `required_features` |
| |
| A bit the reader does not know → `UnsupportedFeature` error (fail-fast; |
| no "read as best we can"). |
| |
| | bit | name | meaning | |
| |-----|----------------|---------| |
| | 0 | `TENSOR_DIR` | binary tensor directory (always set in v2) | |
| | 1 | `BINARY_MASKS` | masks section (§5) present | |
| | 2 | `QUANT_2F` | directory contains `q8_2f`/`vbit` tensors (two-field 𝒲×θ quant) | |
| | 3 | `DELTA_MASKS` | reserved: XOR mask deltas from a parent | |
| | 4 | `HOT_PACKS` | reserved: materialized dense slices | |
| | 5 | `LOOP_MASKS` | mask rows are per VISIT (physical layers × loops, pass-major) — a Looped Transformer's two passes carry independent masks (§5.1) | |
| | 6 | `SKILL_FILE` | the file is a STANDALONE SKILL: a partial tensor set cut against a specific base, bound by `SkillRecord.base_dir_hash` (§9.1). Not runnable — attach with `cortiq skill apply` | |
|
|
| Unknown **header-JSON** fields are ignored (additive evolution); |
| breaking changes go only through feature bits or a `version` bump. |
|
|
| ### 1.2 Validation rules (normative) |
|
|
| The reader MUST return an error (not a default, not a warning) when: |
|
|
| - magic ≠ `CMF\x01` → `InvalidMagic`; |
| - `version` ≠ 2 → `UnsupportedVersion` (v1 is dead: no real v1 files |
| exist, no support program will be started); |
| - an unknown `required_features` bit is set → `UnsupportedFeature`; |
| - any section extends past EOF, `data_off` is not a multiple of 4096, a |
| tensor's `off + nbytes` exceeds `data_len` → `Bounds`; |
| - a tensor name is not UTF-8, dtype is unknown, `ndim > 6` → `Parse`. |
|
|
| Tensor-hash verification is on demand (`cortiq verify`, a loader flag), |
| not on every open: mmap pages are read lazily. |
|
|
| ## 2. Header JSON |
|
|
| UTF-8 JSON, unaligned. Machine-critical data lives in binary sections; |
| JSON carries architecture and provenance — the parts a human reads. |
|
|
| ```jsonc |
| { |
| "format": "cmf", |
| "version": 2, |
| "arch": { |
| "arch_name": "qwen3.5", |
| "hidden_size": 5120, "intermediate_size": 17408, |
| "num_layers": 64, "num_attention_heads": 24, "num_kv_heads": 4, |
| "head_dim": 256, "vocab_size": 248320, |
| "layer_types": ["LinearAttention", "...", "FullAttention"], |
| "rms_norm_eps": 1e-6, |
| "norm_style": "qwen", // "qwen": x̂·w | "gemma": x̂·(1+w) |
| "rope_theta": 1000000.0, |
| "yarn": { // optional global YaRN profile |
| "factor": 128.0, "original_max_position_embeddings": 8192, |
| "beta_fast": 32.0, "beta_slow": 1.0, "attention_factor": 1.485203 |
| }, |
| "attention_heads_per_layer": [48, 72, 72, 72], // optional; length = num_layers |
| "sliding_window": 512, |
| "rope_local_base_freq": 10000.0, |
| "local_partial_rotary_factor": 1.0, |
| "tie_word_embeddings": false, |
| "max_position_embeddings": 262144, |
| "linear_conv_kernel_dim": 4, |
| "linear_num_key_heads": 16, "linear_num_value_heads": 48 |
| }, |
| "quant_type": "Q4_BLOCK", // informational default; truth = per-tensor dtype in the directory |
| "provenance": { "tool": "…", "source_model": "…" } // optional, free-form |
| } |
| ``` |
|
|
| `norm_style` is mandatory for an engine: Gemma-style `(1+w)` applied to |
| Qwen weights is silent garbage across all ~130 normalizations of a |
| forward pass. |
|
|
| Capability dispatch is **tensor-presence driven**: an engine decides |
| per-layer operators by what exists in the directory (q/k biases, |
| qk-norms, output gate by projection width, MoE router, GDN projections) |
| — not by matching model names. New models of a known family load with |
| zero engine changes. |
|
|
| An explicit `SlidingAttention` layer tag selects causal windowed GQA even |
| when the local/global schedule is irregular. Such layers use |
| `sliding_window`, `rope_local_base_freq`, and |
| `local_partial_rotary_factor`; global `FullAttention` layers use |
| `rope_theta`, `partial_rotary_factor`, and optional `yarn`. The optional |
| `attention_heads_per_layer` array overrides the base Q-head count for each |
| layer. Attention projection gating is tensor-presence driven: |
| `self_attn.g_proj.weight [num_heads, hidden]` means per-head |
| `softplus(g_proj·x)` gating immediately before `o_proj` (a |
| `[num_heads·head_dim, hidden]` projection means per-channel gating). |
| These fields and tensor semantics cover Laguna without introducing a |
| model-name-specific execution operator. |
|
|
| ### 2.1 MTP — multi-token prediction (optional) |
|
|
| If the model carries an MTP head (DeepSeek/Qwen style), arch declares: |
|
|
| ```jsonc |
| "mtp": { "num_layers": 1, "share_lm_head": true, "share_embed": true } |
| ``` |
|
|
| MTP tensors are ordinary directory entries under canonical names |
| (`model.mtp.*`): `enorm.weight`, `hnorm.weight`, |
| `eh_proj.weight [hidden, 2·hidden]`, `layers.{i}.*` (a standard |
| transformer block), `norm.weight`. |
|
|
| Semantics: `x = eh_proj·[enorm(embed(t_{p+1})); hnorm(h_p)]` — embedding |
| FIRST (oracle-verified: the reverse order yields exactly 0% acceptance) |
| → block → shared lm_head → draft of token `t_{p+2}`. A reader is not |
| required to execute MTP (metadata + ordinary tensors, additive |
| evolution, no feature bit); the CMF runtime uses the head for |
| speculative decode with a strict guarantee: **output is exactly equal to |
| plain greedy** — a rejected draft is rolled back from KV. |
|
|
| ### 2.2 MoE — mixture-of-experts FFN (optional) |
|
|
| If the model carries MoE layers (Qwen2-MoE / Qwen3-MoE / Qwen3.5-MoE), |
| arch declares: |
|
|
| ```jsonc |
| "moe": { |
| "num_experts": 256, "top_k": 8, "moe_intermediate_size": 512, |
| "norm_topk_prob": true, // Qwen2-MoE: false |
| "shared_expert_intermediate_size": 512, // absent if no shared expert |
| "router_sigmoid": true, // optional; default = softmax |
| "routed_scaling_factor": 2.5 // optional; default = 1 |
| } |
| ``` |
|
|
| Tensors are ordinary directory entries under HF names: |
|
|
| ``` |
| model.layers.{i}.mlp.gate.weight [num_experts, hidden] router |
| model.layers.{i}.mlp.experts.{e}.{gate,up,down}_proj.weight |
| model.layers.{i}.mlp.expert_bias [num_experts] selection only |
| model.layers.{i}.mlp.shared_expert.{gate,up,down}_proj.weight |
| model.layers.{i}.mlp.shared_expert_gate.weight [1, hidden] optional |
| ``` |
|
|
| Which layers are MoE is decided by the PRESENCE of the router in the |
| directory (per-layer, not per-model): Qwen2-MoE's |
| `mlp_only_layers`/`decoder_sparse_step` produce mixed models, and dense |
| layers keep ordinary `mlp.*_proj`. |
|
|
| Execution semantics (HF parity, gated by `tests/moe_parity.sh` across |
| multiple families): by default, softmax over ALL router logits; when |
| `router_sigmoid`, score each expert independently with sigmoid. An optional |
| `expert_bias` affects top-k selection only, not the gathered weights. Select |
| top-k (ties: lower index), optionally renormalize the selected weights, then |
| apply `routed_scaling_factor` and compute Σwₑ·FFNₑ(x). The shared expert is |
| always added: with weight `sigmoid(shared_expert_gate·x)` when that tensor is |
| present, otherwise with weight 1 (Laguna). |
| Experts stay quantized in mmap; per token only the pages of the selected |
| k are touched — the same residency story as skills. Writers SHOULD lay |
| a layer's expert tensors out role-contiguously (all `gate_proj` of |
| experts 0…N−1 back to back, then all `up_proj`, then all `down_proj`) |
| — GPU backends can then treat a layer's expert bank as one region |
| instead of gathering hundreds of slices; the native importer and |
| `moe-defrag` both emit this order. Each expert is a |
| separate directory entry with ITS OWN dtype: that is the carrier of |
| per-expert bit allocation (P15 claim 12) — implemented, gated by |
| `tests/moe_vbit.sh`; the B-field (router selection frequencies via |
| `--route-stats`) was measured end-to-end on a 35B model. |
|
|
| ## 3. Tensor directory |
|
|
| Byte-for-byte the `.vmfc` v2 layout (single canon, shared reference |
| parser): |
|
|
| ``` |
| [0 : 8 ] count : u64 |
| [8 : 16] pool_off : u64 (name-pool offset from section start) |
| [16 : 16 + count·56] 56-byte records: |
| name_off : u32 (relative to pool_off) |
| name_len : u16 |
| dtype : u8 (§3.1) |
| ndim : u8 (≤ 6) |
| shape : u32 × 6 (zero-padded tail) |
| off : u64 (RELATIVE to data_off; multiple of 64) |
| nbytes : u64 |
| hash : u64 (hash64 of the tensor bytes, §8) |
| [pool_off : …] UTF-8 name pool |
| ``` |
|
|
| Tensor names are **1:1 with the source model** |
| (`model.layers.{i}.mlp.gate_proj.weight`, `model.embed_tokens.weight`, |
| `lm_head.weight`, …). The format does not prescribe a tensor set: the |
| directory is the single source of truth for what the blob contains. |
| There is no "computable layout". |
|
|
| ### 3.1 `dtype` |
|
|
| Numbering shared with `.vmfc` (ids are never reused): |
|
|
| | id | name | status in CMF v2 | |
| |----|-----------|------------------| |
| | 0 | `f32` | ✅ read/write | |
| | 1 | `f16` | ✅ read/write (norms and 1-D are always f16) | |
| | 2 | `bf16` | ✅ read/write | |
| | 3 | `q8_row` | ✅ read/write | |
| | 4 | `q4_block`| ✅ read/write | |
| | 5 | `mix8_4` | reserved | |
| | 6 | `u8` | reserved | |
| | 7 | `q4_col` | reserved | |
| | 8 | `vbit` | ✅ read/write (`QUANT_2F` bit), variable 3–8 bit | |
| | 9 | `q8_2f` | ✅ read/write (`QUANT_2F` bit), 𝒲×θ | |
| | 10 | `vbit_ro` | ✅ read/write — `vbit` + in-file row-offset table (O(1) row access); converter default for `--quant vbit` | |
| | 11 | `q4_tiled`| ✅ read/write — q4 in interleaved `[f16 scale][16B nibbles]` tiles (`--quant q4t`) | |
| | 12 | `q1` | ✅ read/write — 1-bit binary, for 1-bit-TRAINED models only (`--quant q1`) | |
| | 13 | `q1s` | ✅ read/write — `q1` base + sparse high-precision outlier overlay (1-bit PTQ of normal checkpoints) | |
| | 14 | `q1t` | ✅ read/write — ternary `{−s, 0, +s}` base-3 tiles + per-row outlier overlay (~2.25 bpw + overlay) | |
| | 15 | `q4tp` | ✅ read/write — `q4_tiled` nibbles with the per-tile scale as a 5-bit rung on a per-row ladder (`--quant q4tp`, or `requant` in place) | |
|
|
| ### 3.2 Quant layouts (canon = `.vmfc`: "quants first, then scales") |
|
|
| - **`q8_row`** (2-D `[out, in]` only): |
| `[int8 : out·in][f16 : out]` — one scale per row, |
| `w = q[o,i]·scale[o]`, `scale[o] = absmax(row_o)/127`. |
| - **`q4_block`**: groups of 32 over the flattened tensor, zero-padded; |
| `[u8 : ceil(n/32)·16][f16 : ceil(n/32)]`. |
| Nibbles: element `2k` low, `2k+1` high; `w = (q − 8)·scale`, |
| `scale = absmax(group)/7`. |
| - **1-D tensors and tensors < 32 elements are always `f16`** |
| (normalization precision at maximal matrix compression). |
| - **`q8_2f`**: `[int8][f16 row-scale][f16 col-field]`, |
| `w = q·scale[o]·col[i]` — the two-field Madelung split 𝒲×θ, validated |
| in vmfcore (+37% at equal size; recovers ~75% of the q8→f16 gap on |
| outlier input channels). |
| - **`vbit`** (2-D only, `in % 32 == 0`; P13 FIG.3): |
| `[u8 bits: rows][f16 scales: rows·in/32][bit-packed rows, MSB-first, |
| each row padded to a byte]`; `w = (u − L)·scale[r,g]`, |
| `L = 2^{b−1}−1`, levels b ∈ {3,4,5,6,8}, floor 3 (claim 13). |
| Allocation b_r: water-filling over the log2 row amplitude toward the |
| tensor's mean budget; for MoE experts the budget is SHARED across the |
| family (layer × projection): the shift `ā_expert − ā_family` is |
| equivalent to joint water-filling over all experts' rows — a loud |
| expert gets more bits, a quiet one is pinned to the floor (P15 |
| claim 12; gate `tests/moe_vbit.sh`). Optionally the allocation takes |
| the product with a B-field — router selection frequencies collected |
| at calibration (`b ∝ log2(A·B)`, truncated Fisher). |
| - **`vbit_ro`** (2-D only, `in % 32 == 0`): the same bits/scales/packed |
| encoding as `vbit`, plus `u32 row_offsets[rows+1]` (relative to the |
| packed area) between the scales and the packed rows — |
| `[u8 bits: rows][f16 scales: rows·in/32][u32 offsets: rows+1][packed]`. |
| Readers get O(1) row access without a prefix scan over bit widths. |
| The byte semantics of `vbit = 8` are untouched; new id on purpose. |
| - **`q4_tiled`** (2-D only, `in % 32 == 0`): |
| `repeat per 32-group { [f16 scale][16B nibbles] }` — 18-byte tiles, |
| one sequential memory stream instead of two distant ones. Values and |
| nibble order are identical to `q4_block`; only the placement of the |
| scale differs (kernel-measured ×1.66 ARM / ×1.13 AVX2 over split). |
| - **`q4tp`** (2-D only, `in % 32 == 0`): |
| `[nibbles: rows·gpr·16][row params: rows × (f16 lo, f16 step)] |
| [codes: rows × ceil(gpr·5/8), 5-bit LSB-first, row-aligned]`, |
| `gpr = in/32`. A tile's scale is `2^(lo[r] + code·step[r])`, so a reader |
| expands one row's 32-rung ladder once and then reads scales by table |
| lookup. Nibble values and order are identical to `q4_tiled`; only the |
| scale's representation differs. 4.17 bits/weight against 4.50 — the |
| f16 scale was 11% of a q4t file. |
| `lo`/`step` come from the row's exact min/max log-scale, so no code is |
| ever out of range and the format needs no escape hatch. Encoders MUST |
| round `lo`/`step` to f16 **before** choosing codes, and MUST quantize the |
| nibbles against the reconstructed scale — otherwise writer and reader |
| disagree, the same trap that makes a q4 encoder round its scale first. |
| - **`q1`** (2-D only, `in % 32 == 0`): |
| `repeat per 32-group { [f16 scale][4B sign bits] }` — 6-byte tiles, |
| 1.5 bits/weight. Bit k of byte j (LSB-first) is weight j·8+k of the |
| group; `w = scale·(2·bit − 1) ∈ {−s, +s}`, `scale = mean|group|` |
| (the L2-optimal binary level). Intended for 1-bit-TRAINED models |
| (Bonsai / BitNet class), where per-group weights already sit on two |
| levels and the encoding is lossless up to f16; as post-training |
| quantization of a normal checkpoint it destroys quality, so |
| converters expose it only as an explicit opt-in. |
| - **`q1s`** (2-D only, `in % 32 == 0`): a `q1` base (identical 6-byte |
| tiles; outliers are EXCLUDED from the group scale) followed by a |
| sparse high-precision overlay: `[u32 count]` then |
| `count × { [u32 flat-index][f16 value] }` — the salient weights kept |
| at full precision (holographic transfer / SpQR-style) and restored |
| verbatim at dequant. Variable length: `expected_nbytes` is |
| undefined, the reader trusts the directory's stored span. Lets a |
| NORMAL checkpoint survive 1-bit where plain `q1` cannot. |
| - **`q1t`** (2-D only, `in % 32 == 0`, `in` must fit `u16`): ternary |
| BitNet-b1.58-style `{−s, 0, +s}`. Base: |
| `repeat per 32-group { [f16 scale][7B base-3 codes] }` — 9-byte |
| tiles, 5 ternary values per byte (3⁵ = 243 ≤ 256; code 0 → 0, |
| 1 → +s, 2 → −s), ~2.25 bits/weight. Then a per-row outlier overlay: |
| `[u32 row_ptr[rows+1]]` followed by `{ [u16 col][f16 value] }` |
| entries grouped by row (row `r`'s outliers are |
| `[row_ptr[r], row_ptr[r+1])`; `col` is a within-row index) — 4 |
| bytes per outlier, no binary search. Capturing the many near-zero |
| weights exactly is the decisive PTQ win over binary. Variable |
| length, same span rule as `q1s`. |
|
|
| ## 4. Weight blob |
|
|
| `data_off` is a multiple of 4096 (page-aligned mmap); every tensor |
| inside starts on a 64-byte boundary (SIMD loads, cache lines). Zero |
| padding between tensors. A reader interprets the blob only through the |
| directory. |
|
|
| ## 5. Masks section |
|
|
| A task mask = bit fields of "what is active" over shared weights |
| (weights do not change — the VMF principle: a skill selects a subset of |
| the condensate). |
|
|
| ``` |
| [0 : 4] n_masks : u32 |
| [4 : 8] meta_len : u32 |
| [8 : 8 + meta_len] JSON meta (§5.1) |
| […] mask blobs, each aligned to 8 from the section start |
| ``` |
|
|
| One mask blob (sizes derived from arch, no internal headers): |
|
|
| ``` |
| [n_layers × ffn_bytes] FFN bitfields ffn_bytes = ceil(intermediate_size / 8) |
| [n_layers × head_bytes] head bitfields head_bytes = ceil(num_attention_heads / 8) |
| [gates_bytes] layer_gates gates_bytes = ceil(num_layers / 8) |
| [n_layers × expert_bytes] expert bitfields OPTIONAL — only when the mask's meta |
| sets "has_expert_fields": true; |
| expert_bytes = ceil(moe.num_experts / 8) |
| ``` |
|
|
| Bit order is LSB-first: neuron `i` = bit `i % 8` of byte `i / 8`; bit |
| set → active. **Tail bits beyond the dimension MUST be zero** (or |
| popcount sees phantom neurons/heads). |
|
|
| The optional expert area (additive: old readers never look past the |
| gates, and each mask's `blob_len` is explicit) makes a task mask narrow |
| MoE ROUTING: bit `e` of layer `l`'s row set → expert `e` is routable |
| for this task; selection then happens over the routable set only, the |
| router softmax renormalizing over it. This is the runtime-switchable |
| twin of §11.1's physical expert defrag — one file with the full expert |
| set serves many specialists (`cortiq moe-mask` writes such masks, |
| `run --task <name>` activates one; verified token-identical to the |
| equivalent runtime restriction). A layer whose row is all-ones is |
| unrestricted; a mask without the area restricts nothing. |
|
|
| ### 5.1 Mask JSON meta |
|
|
| ```jsonc |
| { |
| "default_task": "general", |
| "masks": [{ |
| "task_id": 0, "name": "general", "description": null, |
| "sparsity": 0.62, |
| "quality": { // null = NOT MEASURED (declaring 1.0 is forbidden) |
| "metric": "heldout_ppl_ratio", "value": 0.97, |
| "baseline_dense": 6.10, "n_samples": 512, "dataset_sha256": "…" |
| }, |
| "parent": null, "priority": "Fallback", "has_hot_pack": false, |
| "blob_off": 4096, "blob_len": 139328 // relative to section start |
| }] |
| } |
| ``` |
|
|
| `quality` is a **held-out contract**, not a declaration: a converter |
| without a measured metric writes `null`; the runtime logs a warning when |
| switching to an unmeasured mask. |
|
|
| ## 6. Tokenizer section |
|
|
| The bytes of HuggingFace `tokenizer.json`, verbatim. The model is |
| self-contained: one file = one unit of distribution. A sidecar file |
| remains a debugging fallback. |
|
|
| ### 6.1 Chat bundle (`header.tokenizer_config`) |
| |
| The file — not the runtime binary — defines chat behavior. The header |
| carries an optional block (additive evolution, no feature bit): |
| |
| ```json |
| "tokenizer_config": { |
| "chat_template": "<Jinja template from chat_template.jinja or tokenizer_config.json>", |
| "eos_token_ids": [248044, 248045], |
| "bos_token_id": null, |
| "pad_token_id": 248055 |
| } |
| ``` |
| |
| The runtime renders the template with HF semantics (trim_blocks, |
| lstrip_blocks, loop controls, Python string methods) and stops |
| generation on any id in `eos_token_ids`. Gate: |
| `tests/chat_template_parity.sh` — the runtime render equals reference |
| jinja2 byte-for-byte. Files without the block get a ChatML fallback. |
| |
| ## 7. Sparse index |
| |
| A precomputed bridge "mask → computation skip": active FFN quant groups |
| (32 neurons each) and heads, per (task, layer) pair. |
| |
| > Honest status: the engine takes active indices directly from the mask |
| > bitfields; the index is read and displayed by the CLI but has never |
| > been used in execution. **Deprecation-pending**: writers SHOULD stop |
| > emitting it (readers keep parsing existing files); it is revived only |
| > if the "masks × quantized mmap" path materializes with a measured win. |
| |
| ``` |
| [0 : 4] n_entries : u32 |
| [4 : 8] reserved : u32 (0) |
| entry (4-aligned): |
| task_id : u32 |
| layer_idx : u32 |
| n_groups : u32 |
| n_heads : u32 |
| [u16 × n_groups] active FFN-group indices (sorted) |
| [u8 × n_heads] active head indices (sorted) |
| zero padding to a multiple of 4 |
| ``` |
| |
| A group is active if it contains at least one active mask bit. |
| |
| ## 8. `hash64` |
| |
| A non-cryptographic 64-bit hash of tensor bytes: murmur3 `fmix64` over |
| 64-bit LE words with positional salt `i·0x9E3779B97F4A7C15`, XOR fold, |
| `xor len`, final `fmix64`. Bit-for-bit compatible with |
| `vmfcore.hash64` (Python) and `vmfcore::hash64` (Rust) — hashes of |
| shared tensors match between `.cmf` and `.vmfc` (backbone dedup across |
| skill files is free). |
| |
| Uses: `cortiq verify` (corruption detection), dedup, cache keys. |
| |
| ### 8.1 Section hashes |
| |
| Metadata integrity (not just tensors): |
| |
| - Envelope reserve `[0x70:0x78]` = hash64(header JSON), `[0x78:0x80]` = |
| hash64(directory). Zero = "absent" (older files pass). |
| - The header JSON carries `section_hashes` — hex hash64 of |
| masks/vocab/index (u64 as a JSON number would lose precision past |
| 2^53). The header hash in the envelope transitively covers them. |
| - The envelope itself (first 0x70 bytes) is not hashed: a hash cannot |
| protect itself; corrupted offsets are caught by bounds/hashes further |
| down the chain. |
| - `cortiq verify` checks the whole chain; a single flipped header byte |
| is an error. |
| |
| ### 8.2 Detached signature (authenticity, opt-in) |
| |
| The hash chain proves integrity, not authorship. `cortiq sign` writes a |
| detached `<model>.sig` — JSON `{alg: "ed25519-sha256", pubkey, sha256, |
| sig}`, Ed25519 over the file's SHA-256 — so the container itself is |
| never rewritten and old tooling is untouched. `cortiq verify` checks |
| the signature automatically when the `.sig` sits next to the model; |
| absence is not an error. Key = a 32-byte hex seed file the signer |
| keeps private. |
| |
| ## Anti-features — what the format deliberately does NOT have |
| |
| - **A computable weight layout** — bug class #1 of v1 (writer and reader |
| "computed" the layout independently and diverged). |
| - **Silent fallbacks** — v1 would interpret any garbage file as "a 27B |
| model"; v2 must fail. |
| - **JSON for bit data** — v1 masks in JSON bloated 3–4×. |
| - **Declaration fields** — `quality_score: 1.0` by default, area-law |
| "capacities", Born multipliers in dynamics: a metaphor does not become |
| a format field until it is measured. |
| |
| ## 9. Skills — a swarm in one file (Patent 15, claims 2/12/15) |
| |
| One shared backbone + K per-skill records; no record stores a full |
| model. Storage scales as |backbone| + Σ|deltas|. |
| |
| **Replacement tensors** are ordinary directory entries named |
| `skill.{skill_id}.{name_of_replaced_tensor}`, e.g. |
| `skill.sql.model.layers.3.mlp.gate_proj.weight`. The full logical shape |
| of the replaced tensor (full-shape — NOT low-rank, NOT a diff list, NOT |
| a mask), in any encoding of §3. The per-skill delta index (claim 2) is |
| materialized by the directory: a prefix filter yields skill → |
| byte-offsets; lazy paging = mmap access to exactly those offsets |
| (claim 12). |
| |
| **Registry** — header JSON, additive: |
| |
| ```json |
| "skills": [{ |
| "id": "sql", |
| "name": "SQL assistant", |
| "layers": [3, 4, 5], |
| "selection": {"metric": "mse", "phi_layer": 20, |
| "mean": "<f16 base64>", "basis": "<f16 base64>"}, |
| "input_mask_task": null, |
| "quality": {"metric": "ppl", "backbone": 21.4, "overlaid": 17.9, |
| "dataset_sha256": "…"} |
| }] |
| ``` |
| |
| `selection` holds the affine-subspace parameters for recon-argmin |
| routing (`E = ‖r − BBᵀr‖²/‖φ‖²`, choose the skill with minimal E); the |
| file is self-sufficient for selection. `quality` is the honest claim-16 |
| contract (overlaid vs backbone on held-out data). |
| |
| **Execution semantics (claims 1/3/18)**: tensor-source indirection — for |
| every tensor the runtime reads EITHER the backbone entry OR |
| `skill.{active}.{name}` if present; replacement instead of addition, a |
| full per-skill model is never assembled (all tensors are pointers into |
| one mmap). Soft superposition (claim 14): blended working tensors |
| `Σwᵢ·Tᵢ`, `wᵢ = softmax(−E/T)`. |
| |
| **Append-only growth (claim 11)**: adding a skill = appending new |
| tensors at the file tail + re-emitting directory/header/index at the |
| tail + updating envelope offsets in place (offset 0 is fixed). Bytes and |
| offsets of previously written tensors never change; old dir/header bytes |
| become dead section tails (compatible: readers navigate only through the |
| envelope). Compaction (`converter/cmf_compact.py`) = a plain rewrite. |
| |
| ### 9.1 Standalone skill files (`SKILL_FILE`, bit 6) |
| |
| A skill can also travel WITHOUT its backbone: a `.cmf` whose tensor set |
| is only what a bake changed (plus the mask catalog), bound to the base |
| it was cut against by identity keys in the registry record: |
| |
| ```json |
| "skills": [{ |
| "id": "gfx-html", |
| "layers": [0, 1, "...", 21], |
| "base_dir_hash": "9f22593eb458bc6f", |
| "base_arch": "nanbeige", |
| "task": "specialist", |
| "provenance": {"corpus": "…", "tensors": 30} |
| }] |
| ``` |
| |
| - `base_dir_hash` — hex `hash64` of the BASE file's tensor-directory |
| bytes (the same value the envelope carries at `[0x78]`). A skill is a |
| delta against exact bytes, not against an architecture: `apply` MUST |
| refuse a base whose directory hash differs (an explicit `--force` |
| may override; the result is out of spec). |
| - `base_arch`, `task`, `provenance` — informative keys: the human check, |
| the mask-catalog task the skill activates, and where it came from. |
|
|
| Any record with `base_dir_hash` present raises feature bit 6, so a |
| pre-bit reader refuses the file loudly and a runtime that knows the bit |
| refuses to RUN it (a partial tensor set is not a model) and points to |
| `cortiq skill apply <base> <skill> -o out.cmf`, which verifies the key, |
| overlays tensors and masks over the base, and writes a complete file — |
| byte-equivalent to the specialist the skill was cut from. |
|
|
| Lifecycle: `skill bake` (specialist) → `skill export --base` (delta + |
| keys) → publish the small file → `skill apply` on any copy of the base. |
|
|
| Status: fully implemented and gated (container + indirection, |
| production recipes, recon-argmin routing, append-only + compaction, |
| soft-blend); claim 16 met by measurement (−24.9% task-PPL in the |
| runtime). |
|
|
| ## 10. Sharding — a model in N files |
|
|
| Naming: `{base}-{no:05}-of-{count:05}.cmf` (spiritually compatible with |
| safetensors). The user opens ANY name; the runtime normalizes to shard 1 |
| and picks up siblings by pattern. |
|
|
| **Every shard is a standalone valid .cmf**: full envelope, header JSON, |
| a directory of ITS OWN tensors, its own data blob, its own hashes |
| (`section_hashes` + per-tensor). `cortiq verify` works on any single |
| shard without its siblings. |
|
|
| Each shard's header carries: |
|
|
| ```json |
| "shard": { "no": 1, "count": 5 } |
| ``` |
|
|
| No block = an ordinary single file (backward compatible: old readers see |
| shard 1 as a valid but incomplete model and fail honestly on the missing |
| tensor). |
|
|
| **Content distribution**: tensors are split greedily in canonical order |
| (`--shard-max-gb` threshold, rough f32 size); the masks/vocab/sparse |
| index sections, `tokenizer_config` (chat bundle) and the `skills` |
| registry live ONLY in shard 1 — the rest have empty sections and |
| `tokenizer_config: null`. Skill tensors (`skill.{id}.*`) are distributed |
| as ordinary directory entries — the shard-1 registry references them by |
| name through the merged directory. |
|
|
| **Loading** (`CmfModel::open_sharded`): open shard 1 → mmap all siblings |
| → merge directories (each entry remembers its shard index — a runtime |
| field, never written to disk) → the runtime then works as with a single |
| file. Errors: opening a non-first shard directly, a missing sibling, a |
| `count` mismatch. |
|
|
| Gate (Qwen3.5-0.8B q8_2f, 5 shards ≤ 0.6 GB): sharded PPL == unsharded |
| byte-exactly on the same binary; `verify` green on every shard alone. |
| |
| ## 11. Defragmentation — physical pruning (USPTO App. 19/452,464, claims 9/10 — [PATENTS.md](../PATENTS.md)) |
| |
| A mask (§5) is **virtual sparsity**: pruned neurons are flagged but still |
| stored in full (all tasks share one backbone — you cannot physically cut |
| it until you commit to ONE task). Defragmentation turns virtual sparsity |
| into **physical compression**: pruned FFN neurons are dropped from the |
| file — they are **neither stored nor computed**. This is Factory-Hard → |
| defrag from the DTG-MA application (19/452,464): "bake one mask into the weights" and emit a |
| standalone compact `.cmf`. |
| |
| **Representation — no new feature bit, backward compatible.** Physical |
| pruning is expressed ONLY by smaller tensor shapes in the directory (§3 |
| "no computable layout"; the directory is the sole shape authority). The |
| runtime derives the FFN size from the tensor shape (`gate_proj.rows()`), |
| not from `arch.intermediate_size`, so a defragged file is an ordinary |
| smaller dense model that existing readers load unchanged. The masks |
| section (§5) is **absent** in a defragged file (the mask is the identity |
| after pruning). `arch.intermediate_size` becomes nominal (= the per-layer |
| max); the true size lives in each tensor. |
|
|
| **Per-layer variance — better than the patent.** Because the directory |
| carries an arbitrary shape per tensor, each layer shrinks to its OWN |
| live-neuron count. The patent must truncate every layer to `max(active)` |
| (one bottleneck layer caps the ratio at 80.2% vs. 94% achievable) — CMF |
| has no such limit. |
|
|
| **Invariants (mandatory):** |
|
|
| - per-layer triple: `gate_proj.rows() == up_proj.rows() == |
| down_proj.cols() == inter'ₗ`, and `down_proj.rows() == hidden_size`. |
| One keep-set indexes all three (gate row i, up row i, down col i are |
| the same neuron); |
| - neuron axis: rows (axis 0) for `gate_proj`/`up_proj`, columns (axis 1) |
| for `down_proj`; |
| - quant group of 32: the `down_proj` neuron axis is its COLUMNS, and |
| `vbit`/`q4_block` require `in % 32 == 0`. A `down_proj` whose `inter'` |
| is not a multiple of 32 is written as `q8_2f` (per-row scale — no |
| column constraint; the converter downgrades automatically). `gate/up` |
| drop rows, so their columns (= hidden) are unaffected; |
| - NOT a byte truncation: quant scales are per-group/per-row, so pruning |
| is dequant → gather live neurons → **requant** at the smaller shape |
| (the `q8_2f` col-field / `vbit` scales of `down_proj` regenerate for |
| the shrunk column set); tensor hashes are recomputed; |
| - `hidden_size`, `embed_tokens`, `lm_head`, and norms are untouched |
| (skill-selection subspaces depend on hidden). |
|
|
| **One task, standalone file.** Defrag is destructive: one `.cmf` bakes |
| exactly one task. Multi-task serving stays on masks (§5) or per-skill |
| replacement tensors (§9). |
|
|
| **Provenance (honest contract).** Header `provenance.defrag`: |
|
|
| ```jsonc |
| "defrag": { |
| "source_skill": "…/skill_ru", |
| "pre_intermediate": 3072, |
| "post_intermediate_max": 640, |
| "kept_per_layer": [608, 640, 512, ...], |
| "pruned_ratio": 0.803 |
| } |
| ``` |
|
|
| Numerically the dense output of a defragged model is IDENTICAL to the |
| masked output before quantization (a dead neuron contributes `act·0` |
| under a mask and is simply absent after defrag); after quantization the |
| only difference comes from quantizing the smaller matrices. |
|
|
| **Scope:** dense FFN neurons here; MoE experts in §11.1. Attention-head |
| pruning (the head count is a global runtime scalar) is out of scope. |
|
|
| ### 11.1 MoE expert defrag (`cortiq moe-defrag`) |
|
|
| The MoE twin of §11, driven by the routing B-field instead of a neuron |
| mask: expert usage is strongly task-conditional (measured on a 34.7B |
| coder: the top-64 expert sets for code vs prose overlap with Jaccard |
| 0.25), so a one-task file can drop the experts that task never routes |
| to. From a `CMF_MOE_STATS` dump (per-layer expert-selection counts over |
| a task-representative run), keep per layer the smallest top expert set |
| reaching `--cover` of the recorded routing mass; drop the rest. |
|
|
| **Representation — same philosophy as §11, no feature bit.** |
|
|
| - Kept experts are renumbered into a CONTIGUOUS per-layer prefix |
| `mlp.experts.0 … mlp.experts.{k−1}` preserving relative order; a |
| reader enumerates a layer's experts by tensor PRESENCE up to |
| `arch.moe.num_experts`, which becomes nominal (= the original |
| count) — mirroring §11's rule for `intermediate_size`. |
| - The router tensor's rows are gathered to match, in the same order: |
| `mlp.gate.weight` becomes `[kept_l, hidden]`, and |
| `router.rows() == (number of expert entries present)` is a load-time |
| invariant. `top_k` clamps to the per-layer expert count. |
| - Selection semantics are unchanged (§2.2): the softmax simply |
| renormalizes over the kept set. The identical restriction can be |
| applied at RUNTIME without rewriting the file |
| (`CMF_MOE_MASK=<stats.json>` + `CMF_MOE_MASK_COVER`) — the two are |
| mathematically equal, which is how a cover level is perplexity-gated |
| before committing to the cut. |
| - Expert payloads are copied verbatim (no requant — the expert axis is |
| whole tensors, not quant groups), so the surviving weights are |
| byte-identical to the source and the rewrite streams from the source |
| mmap. |
|
|
| Measured reference (KAT-Coder 34.7B-A3B, code-calibrated, cover 0.95): |
| 19.6 → 12.7 GB (−35%), held-out code perplexity +2.8%, and on a 24 GB |
| machine — where the full model paged — decode ×1.8, prefill ×3.3. |
|
|
| Off-task quality degrades by design; like §11, one defragged file bakes |
| one task. Multi-task serving stays on the full expert set. |
|
|
| **Producing it (native Rust):** |
|
|
| ``` |
| cortiq convert --model <hf_dir_or_repo> --defrag <skill_dir> \ |
| --quant q8_2f --output model.cmf |
| ``` |
|
|
| `<skill_dir>` carries baked FFN overlays (`tensors/*.npy`) and, if |
| available, a keep-set `ffn_keep.npy` (bool `[n_layers, intermediate]`, |
| True = live) from the pruning pipeline. Without `ffn_keep.npy` the |
| keep-set is autodetected from all-zero `down_proj` columns (the |
| Factory-Hard bake). The mask-training / bake step lives in the private |
| research pipeline; the public tool only consumes its artifacts. |
|
|
| ## 12. Pipeline containers — text-to-image in one file |
|
|
| The same envelope/directory/blob machinery carries non-LLM pipelines. |
| The only differences are the `arch_name` tag and namespaced tensor |
| names; no new sections, no feature bit (a reader that does not execute |
| the pipeline still validates and inspects the file). |
|
|
| Current instance — `arch_name: "lumina2-image"` (Lumina-Image 2.0, |
| `cortiq imagine-pack` / `cortiq imagine`): one file packs the whole |
| text-to-image stack. |
|
|
| - **Namespaces**: `te.*` — the text-encoder transformer (a Gemma-2 |
| class LLM; the header's `arch` block describes THIS component, so |
| generic tooling reads meaningful dimensions), `dit.*` — the Next-DiT |
| denoiser, `vae.*` — the VAE decoder. Component config JSONs ride as |
| `{prefix}.config_json` u8 tensors — the file is self-sufficient. |
| - **Quantization**: per-tensor as always (§3 directory is the truth) — |
| typically q4t/q8 matrices for te/dit, f16 for VAE convolutions and |
| norms. |
| - **Tokenizer section** (§6) carries the text encoder's tokenizer; |
| `provenance.pipeline` + `provenance.components` name the recipe. |
|
|
| The measured reference lives in the README (512 px on CPU in minutes, |
| Metal whole-DiT-block graph on Apple silicon; the wgpu path serves |
| discrete cards and phones). |
|
|
| --- |
|
|
| *Related: [COMPARISON.md](COMPARISON.md) (CMF vs. other model formats), |
| [project README](../README.md) (overview and quick start), |
| `python/cmf_reader.py` (standalone reader: stdlib + numpy, reads every |
| dtype, shards, skills, verify).* |
|
|