File size: 37,682 Bytes
7d1676b e9433fd 7d1676b 1a0c6c0 7d1676b 1a0c6c0 7d1676b 1a0c6c0 7d1676b 1a0c6c0 7d1676b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 | # CMF v2 — Format Specification
*Languages: **English** · [Русский](SPEC.ru.md) · [中文](SPEC.zh.md)*
**Cortiq Model Format** — a single file carrying everything needed for
sparse, task-routed inference: quantized weights, tokenizer, per-task
masks, a precomputed sparse index — and, uniquely, a **swarm of skills**
sharing one backbone (Patent 15).
> Normative source: this document. Reference
> implementations: Rust reader/runtime (`crates/cortiq-core`,
> `crates/cortiq-engine`), Python writer (`converter/`), and a
> standalone Python reader (`python/cmf_reader.py`, stdlib + numpy).
Three requirements, in priority order:
1. **Correct.** No silent corruption modes: strict magic, version,
`required_features`, bounds on every section, a 64-bit hash for every
tensor. A file is either valid or open() returns an error — there is
no third state.
2. **Fast.** The weight section is page-aligned for mmap, every tensor
is 64-byte aligned (zero-copy SIMD), the tensor directory is binary —
read without parsing. Cold (masked-out) weights cost no RSS.
3. **Compact.** Masks are bit-packed (1 bit per neuron), weights are
q4/q8/variable-bit, the whole file is addressed by one 128-byte
envelope.
Harmony comes not from feature count but from **a single canon**: one
layout per level (envelope, directory, quant block, mask), byte-for-byte
compatible with the validated `.vmfc` v2 format where the domains
overlap (tensor directory, quant layouts, `hash64`). Never two
definitions of the same thing.
Physical basis (VMF): the model is a vacuum condensate 𝒲; a skill is its
regular core above a critical density; a task mask selects an active
subset without changing weights. The format carries the consequences of
that physics (two-field 𝒲×θ quantization, Born importance, critical mask
threshold) — but **only those confirmed by measurement**.
---
## 1. Envelope (fixed 128 bytes)
All integers are little-endian.
```
[0x00 : 0x04] magic = b"CMF\x01" (4 bytes)
[0x04 : 0x08] version : u32 = 2
[0x08 : 0x0C] flags : u32 (reserved, 0)
[0x0C : 0x10] required_features : u32 (bitmask, §1.1)
[0x10 : 0x18] header_off : u64 (= 128)
[0x18 : 0x20] header_len : u64 — JSON header (§2)
[0x20 : 0x28] dir_off : u64 — tensor directory (§3)
[0x28 : 0x30] dir_len : u64
[0x30 : 0x38] data_off : u64 — weight blob; multiple of 4096 (§4)
[0x38 : 0x40] data_len : u64
[0x40 : 0x48] masks_off : u64 — masks section (§5); 0 = absent
[0x48 : 0x50] masks_len : u64
[0x50 : 0x58] vocab_off : u64 — tokenizer (§6); 0 = absent
[0x58 : 0x60] vocab_len : u64
[0x60 : 0x68] index_off : u64 — sparse index (§7); 0 = absent
[0x68 : 0x70] index_len : u64
[0x70 : 0x80] reserved : 16 bytes (§8.1: header/dir hashes)
```
Section order on disk: envelope → header JSON → directory → **weight
blob (aligned to 4096)** → masks → vocab → sparse index. A reader MUST
address sections ONLY through the envelope, never by assuming order.
### 1.1 `required_features`
A bit the reader does not know → `UnsupportedFeature` error (fail-fast;
no "read as best we can").
| bit | name | meaning |
|-----|----------------|---------|
| 0 | `TENSOR_DIR` | binary tensor directory (always set in v2) |
| 1 | `BINARY_MASKS` | masks section (§5) present |
| 2 | `QUANT_2F` | directory contains `q8_2f`/`vbit` tensors (two-field 𝒲×θ quant) |
| 3 | `DELTA_MASKS` | reserved: XOR mask deltas from a parent |
| 4 | `HOT_PACKS` | reserved: materialized dense slices |
| 5 | `LOOP_MASKS` | mask rows are per VISIT (physical layers × loops, pass-major) — a Looped Transformer's two passes carry independent masks (§5.1) |
| 6 | `SKILL_FILE` | the file is a STANDALONE SKILL: a partial tensor set cut against a specific base, bound by `SkillRecord.base_dir_hash` (§9.1). Not runnable — attach with `cortiq skill apply` |
Unknown **header-JSON** fields are ignored (additive evolution);
breaking changes go only through feature bits or a `version` bump.
### 1.2 Validation rules (normative)
The reader MUST return an error (not a default, not a warning) when:
- magic ≠ `CMF\x01` → `InvalidMagic`;
- `version` ≠ 2 → `UnsupportedVersion` (v1 is dead: no real v1 files
exist, no support program will be started);
- an unknown `required_features` bit is set → `UnsupportedFeature`;
- any section extends past EOF, `data_off` is not a multiple of 4096, a
tensor's `off + nbytes` exceeds `data_len` → `Bounds`;
- a tensor name is not UTF-8, dtype is unknown, `ndim > 6` → `Parse`.
Tensor-hash verification is on demand (`cortiq verify`, a loader flag),
not on every open: mmap pages are read lazily.
## 2. Header JSON
UTF-8 JSON, unaligned. Machine-critical data lives in binary sections;
JSON carries architecture and provenance — the parts a human reads.
```jsonc
{
"format": "cmf",
"version": 2,
"arch": {
"arch_name": "qwen3.5",
"hidden_size": 5120, "intermediate_size": 17408,
"num_layers": 64, "num_attention_heads": 24, "num_kv_heads": 4,
"head_dim": 256, "vocab_size": 248320,
"layer_types": ["LinearAttention", "...", "FullAttention"],
"rms_norm_eps": 1e-6,
"norm_style": "qwen", // "qwen": x̂·w | "gemma": x̂·(1+w)
"rope_theta": 1000000.0,
"yarn": { // optional global YaRN profile
"factor": 128.0, "original_max_position_embeddings": 8192,
"beta_fast": 32.0, "beta_slow": 1.0, "attention_factor": 1.485203
},
"attention_heads_per_layer": [48, 72, 72, 72], // optional; length = num_layers
"sliding_window": 512,
"rope_local_base_freq": 10000.0,
"local_partial_rotary_factor": 1.0,
"tie_word_embeddings": false,
"max_position_embeddings": 262144,
"linear_conv_kernel_dim": 4,
"linear_num_key_heads": 16, "linear_num_value_heads": 48
},
"quant_type": "Q4_BLOCK", // informational default; truth = per-tensor dtype in the directory
"provenance": { "tool": "…", "source_model": "…" } // optional, free-form
}
```
`norm_style` is mandatory for an engine: Gemma-style `(1+w)` applied to
Qwen weights is silent garbage across all ~130 normalizations of a
forward pass.
Capability dispatch is **tensor-presence driven**: an engine decides
per-layer operators by what exists in the directory (q/k biases,
qk-norms, output gate by projection width, MoE router, GDN projections)
— not by matching model names. New models of a known family load with
zero engine changes.
An explicit `SlidingAttention` layer tag selects causal windowed GQA even
when the local/global schedule is irregular. Such layers use
`sliding_window`, `rope_local_base_freq`, and
`local_partial_rotary_factor`; global `FullAttention` layers use
`rope_theta`, `partial_rotary_factor`, and optional `yarn`. The optional
`attention_heads_per_layer` array overrides the base Q-head count for each
layer. Attention projection gating is tensor-presence driven:
`self_attn.g_proj.weight [num_heads, hidden]` means per-head
`softplus(g_proj·x)` gating immediately before `o_proj` (a
`[num_heads·head_dim, hidden]` projection means per-channel gating).
These fields and tensor semantics cover Laguna without introducing a
model-name-specific execution operator.
### 2.1 MTP — multi-token prediction (optional)
If the model carries an MTP head (DeepSeek/Qwen style), arch declares:
```jsonc
"mtp": { "num_layers": 1, "share_lm_head": true, "share_embed": true }
```
MTP tensors are ordinary directory entries under canonical names
(`model.mtp.*`): `enorm.weight`, `hnorm.weight`,
`eh_proj.weight [hidden, 2·hidden]`, `layers.{i}.*` (a standard
transformer block), `norm.weight`.
Semantics: `x = eh_proj·[enorm(embed(t_{p+1})); hnorm(h_p)]` — embedding
FIRST (oracle-verified: the reverse order yields exactly 0% acceptance)
→ block → shared lm_head → draft of token `t_{p+2}`. A reader is not
required to execute MTP (metadata + ordinary tensors, additive
evolution, no feature bit); the CMF runtime uses the head for
speculative decode with a strict guarantee: **output is exactly equal to
plain greedy** — a rejected draft is rolled back from KV.
### 2.2 MoE — mixture-of-experts FFN (optional)
If the model carries MoE layers (Qwen2-MoE / Qwen3-MoE / Qwen3.5-MoE),
arch declares:
```jsonc
"moe": {
"num_experts": 256, "top_k": 8, "moe_intermediate_size": 512,
"norm_topk_prob": true, // Qwen2-MoE: false
"shared_expert_intermediate_size": 512, // absent if no shared expert
"router_sigmoid": true, // optional; default = softmax
"routed_scaling_factor": 2.5 // optional; default = 1
}
```
Tensors are ordinary directory entries under HF names:
```
model.layers.{i}.mlp.gate.weight [num_experts, hidden] router
model.layers.{i}.mlp.experts.{e}.{gate,up,down}_proj.weight
model.layers.{i}.mlp.expert_bias [num_experts] selection only
model.layers.{i}.mlp.shared_expert.{gate,up,down}_proj.weight
model.layers.{i}.mlp.shared_expert_gate.weight [1, hidden] optional
```
Which layers are MoE is decided by the PRESENCE of the router in the
directory (per-layer, not per-model): Qwen2-MoE's
`mlp_only_layers`/`decoder_sparse_step` produce mixed models, and dense
layers keep ordinary `mlp.*_proj`.
Execution semantics (HF parity, gated by `tests/moe_parity.sh` across
multiple families): by default, softmax over ALL router logits; when
`router_sigmoid`, score each expert independently with sigmoid. An optional
`expert_bias` affects top-k selection only, not the gathered weights. Select
top-k (ties: lower index), optionally renormalize the selected weights, then
apply `routed_scaling_factor` and compute Σwₑ·FFNₑ(x). The shared expert is
always added: with weight `sigmoid(shared_expert_gate·x)` when that tensor is
present, otherwise with weight 1 (Laguna).
Experts stay quantized in mmap; per token only the pages of the selected
k are touched — the same residency story as skills. Writers SHOULD lay
a layer's expert tensors out role-contiguously (all `gate_proj` of
experts 0…N−1 back to back, then all `up_proj`, then all `down_proj`)
— GPU backends can then treat a layer's expert bank as one region
instead of gathering hundreds of slices; the native importer and
`moe-defrag` both emit this order. Each expert is a
separate directory entry with ITS OWN dtype: that is the carrier of
per-expert bit allocation (P15 claim 12) — implemented, gated by
`tests/moe_vbit.sh`; the B-field (router selection frequencies via
`--route-stats`) was measured end-to-end on a 35B model.
## 3. Tensor directory
Byte-for-byte the `.vmfc` v2 layout (single canon, shared reference
parser):
```
[0 : 8 ] count : u64
[8 : 16] pool_off : u64 (name-pool offset from section start)
[16 : 16 + count·56] 56-byte records:
name_off : u32 (relative to pool_off)
name_len : u16
dtype : u8 (§3.1)
ndim : u8 (≤ 6)
shape : u32 × 6 (zero-padded tail)
off : u64 (RELATIVE to data_off; multiple of 64)
nbytes : u64
hash : u64 (hash64 of the tensor bytes, §8)
[pool_off : …] UTF-8 name pool
```
Tensor names are **1:1 with the source model**
(`model.layers.{i}.mlp.gate_proj.weight`, `model.embed_tokens.weight`,
`lm_head.weight`, …). The format does not prescribe a tensor set: the
directory is the single source of truth for what the blob contains.
There is no "computable layout".
### 3.1 `dtype`
Numbering shared with `.vmfc` (ids are never reused):
| id | name | status in CMF v2 |
|----|-----------|------------------|
| 0 | `f32` | ✅ read/write |
| 1 | `f16` | ✅ read/write (norms and 1-D are always f16) |
| 2 | `bf16` | ✅ read/write |
| 3 | `q8_row` | ✅ read/write |
| 4 | `q4_block`| ✅ read/write |
| 5 | `mix8_4` | reserved |
| 6 | `u8` | reserved |
| 7 | `q4_col` | reserved |
| 8 | `vbit` | ✅ read/write (`QUANT_2F` bit), variable 3–8 bit |
| 9 | `q8_2f` | ✅ read/write (`QUANT_2F` bit), 𝒲×θ |
| 10 | `vbit_ro` | ✅ read/write — `vbit` + in-file row-offset table (O(1) row access); converter default for `--quant vbit` |
| 11 | `q4_tiled`| ✅ read/write — q4 in interleaved `[f16 scale][16B nibbles]` tiles (`--quant q4t`) |
| 12 | `q1` | ✅ read/write — 1-bit binary, for 1-bit-TRAINED models only (`--quant q1`) |
| 13 | `q1s` | ✅ read/write — `q1` base + sparse high-precision outlier overlay (1-bit PTQ of normal checkpoints) |
| 14 | `q1t` | ✅ read/write — ternary `{−s, 0, +s}` base-3 tiles + per-row outlier overlay (~2.25 bpw + overlay) |
| 15 | `q4tp` | ✅ read/write — `q4_tiled` nibbles with the per-tile scale as a 5-bit rung on a per-row ladder (`--quant q4tp`, or `requant` in place) |
### 3.2 Quant layouts (canon = `.vmfc`: "quants first, then scales")
- **`q8_row`** (2-D `[out, in]` only):
`[int8 : out·in][f16 : out]` — one scale per row,
`w = q[o,i]·scale[o]`, `scale[o] = absmax(row_o)/127`.
- **`q4_block`**: groups of 32 over the flattened tensor, zero-padded;
`[u8 : ceil(n/32)·16][f16 : ceil(n/32)]`.
Nibbles: element `2k` low, `2k+1` high; `w = (q − 8)·scale`,
`scale = absmax(group)/7`.
- **1-D tensors and tensors < 32 elements are always `f16`**
(normalization precision at maximal matrix compression).
- **`q8_2f`**: `[int8][f16 row-scale][f16 col-field]`,
`w = q·scale[o]·col[i]` — the two-field Madelung split 𝒲×θ, validated
in vmfcore (+37% at equal size; recovers ~75% of the q8→f16 gap on
outlier input channels).
- **`vbit`** (2-D only, `in % 32 == 0`; P13 FIG.3):
`[u8 bits: rows][f16 scales: rows·in/32][bit-packed rows, MSB-first,
each row padded to a byte]`; `w = (u − L)·scale[r,g]`,
`L = 2^{b−1}−1`, levels b ∈ {3,4,5,6,8}, floor 3 (claim 13).
Allocation b_r: water-filling over the log2 row amplitude toward the
tensor's mean budget; for MoE experts the budget is SHARED across the
family (layer × projection): the shift `ā_expert − ā_family` is
equivalent to joint water-filling over all experts' rows — a loud
expert gets more bits, a quiet one is pinned to the floor (P15
claim 12; gate `tests/moe_vbit.sh`). Optionally the allocation takes
the product with a B-field — router selection frequencies collected
at calibration (`b ∝ log2(A·B)`, truncated Fisher).
- **`vbit_ro`** (2-D only, `in % 32 == 0`): the same bits/scales/packed
encoding as `vbit`, plus `u32 row_offsets[rows+1]` (relative to the
packed area) between the scales and the packed rows —
`[u8 bits: rows][f16 scales: rows·in/32][u32 offsets: rows+1][packed]`.
Readers get O(1) row access without a prefix scan over bit widths.
The byte semantics of `vbit = 8` are untouched; new id on purpose.
- **`q4_tiled`** (2-D only, `in % 32 == 0`):
`repeat per 32-group { [f16 scale][16B nibbles] }` — 18-byte tiles,
one sequential memory stream instead of two distant ones. Values and
nibble order are identical to `q4_block`; only the placement of the
scale differs (kernel-measured ×1.66 ARM / ×1.13 AVX2 over split).
- **`q4tp`** (2-D only, `in % 32 == 0`):
`[nibbles: rows·gpr·16][row params: rows × (f16 lo, f16 step)]
[codes: rows × ceil(gpr·5/8), 5-bit LSB-first, row-aligned]`,
`gpr = in/32`. A tile's scale is `2^(lo[r] + code·step[r])`, so a reader
expands one row's 32-rung ladder once and then reads scales by table
lookup. Nibble values and order are identical to `q4_tiled`; only the
scale's representation differs. 4.17 bits/weight against 4.50 — the
f16 scale was 11% of a q4t file.
`lo`/`step` come from the row's exact min/max log-scale, so no code is
ever out of range and the format needs no escape hatch. Encoders MUST
round `lo`/`step` to f16 **before** choosing codes, and MUST quantize the
nibbles against the reconstructed scale — otherwise writer and reader
disagree, the same trap that makes a q4 encoder round its scale first.
- **`q1`** (2-D only, `in % 32 == 0`):
`repeat per 32-group { [f16 scale][4B sign bits] }` — 6-byte tiles,
1.5 bits/weight. Bit k of byte j (LSB-first) is weight j·8+k of the
group; `w = scale·(2·bit − 1) ∈ {−s, +s}`, `scale = mean|group|`
(the L2-optimal binary level). Intended for 1-bit-TRAINED models
(Bonsai / BitNet class), where per-group weights already sit on two
levels and the encoding is lossless up to f16; as post-training
quantization of a normal checkpoint it destroys quality, so
converters expose it only as an explicit opt-in.
- **`q1s`** (2-D only, `in % 32 == 0`): a `q1` base (identical 6-byte
tiles; outliers are EXCLUDED from the group scale) followed by a
sparse high-precision overlay: `[u32 count]` then
`count × { [u32 flat-index][f16 value] }` — the salient weights kept
at full precision (holographic transfer / SpQR-style) and restored
verbatim at dequant. Variable length: `expected_nbytes` is
undefined, the reader trusts the directory's stored span. Lets a
NORMAL checkpoint survive 1-bit where plain `q1` cannot.
- **`q1t`** (2-D only, `in % 32 == 0`, `in` must fit `u16`): ternary
BitNet-b1.58-style `{−s, 0, +s}`. Base:
`repeat per 32-group { [f16 scale][7B base-3 codes] }` — 9-byte
tiles, 5 ternary values per byte (3⁵ = 243 ≤ 256; code 0 → 0,
1 → +s, 2 → −s), ~2.25 bits/weight. Then a per-row outlier overlay:
`[u32 row_ptr[rows+1]]` followed by `{ [u16 col][f16 value] }`
entries grouped by row (row `r`'s outliers are
`[row_ptr[r], row_ptr[r+1])`; `col` is a within-row index) — 4
bytes per outlier, no binary search. Capturing the many near-zero
weights exactly is the decisive PTQ win over binary. Variable
length, same span rule as `q1s`.
## 4. Weight blob
`data_off` is a multiple of 4096 (page-aligned mmap); every tensor
inside starts on a 64-byte boundary (SIMD loads, cache lines). Zero
padding between tensors. A reader interprets the blob only through the
directory.
## 5. Masks section
A task mask = bit fields of "what is active" over shared weights
(weights do not change — the VMF principle: a skill selects a subset of
the condensate).
```
[0 : 4] n_masks : u32
[4 : 8] meta_len : u32
[8 : 8 + meta_len] JSON meta (§5.1)
[…] mask blobs, each aligned to 8 from the section start
```
One mask blob (sizes derived from arch, no internal headers):
```
[n_layers × ffn_bytes] FFN bitfields ffn_bytes = ceil(intermediate_size / 8)
[n_layers × head_bytes] head bitfields head_bytes = ceil(num_attention_heads / 8)
[gates_bytes] layer_gates gates_bytes = ceil(num_layers / 8)
[n_layers × expert_bytes] expert bitfields OPTIONAL — only when the mask's meta
sets "has_expert_fields": true;
expert_bytes = ceil(moe.num_experts / 8)
```
Bit order is LSB-first: neuron `i` = bit `i % 8` of byte `i / 8`; bit
set → active. **Tail bits beyond the dimension MUST be zero** (or
popcount sees phantom neurons/heads).
The optional expert area (additive: old readers never look past the
gates, and each mask's `blob_len` is explicit) makes a task mask narrow
MoE ROUTING: bit `e` of layer `l`'s row set → expert `e` is routable
for this task; selection then happens over the routable set only, the
router softmax renormalizing over it. This is the runtime-switchable
twin of §11.1's physical expert defrag — one file with the full expert
set serves many specialists (`cortiq moe-mask` writes such masks,
`run --task <name>` activates one; verified token-identical to the
equivalent runtime restriction). A layer whose row is all-ones is
unrestricted; a mask without the area restricts nothing.
### 5.1 Mask JSON meta
```jsonc
{
"default_task": "general",
"masks": [{
"task_id": 0, "name": "general", "description": null,
"sparsity": 0.62,
"quality": { // null = NOT MEASURED (declaring 1.0 is forbidden)
"metric": "heldout_ppl_ratio", "value": 0.97,
"baseline_dense": 6.10, "n_samples": 512, "dataset_sha256": "…"
},
"parent": null, "priority": "Fallback", "has_hot_pack": false,
"blob_off": 4096, "blob_len": 139328 // relative to section start
}]
}
```
`quality` is a **held-out contract**, not a declaration: a converter
without a measured metric writes `null`; the runtime logs a warning when
switching to an unmeasured mask.
## 6. Tokenizer section
The bytes of HuggingFace `tokenizer.json`, verbatim. The model is
self-contained: one file = one unit of distribution. A sidecar file
remains a debugging fallback.
### 6.1 Chat bundle (`header.tokenizer_config`)
The file — not the runtime binary — defines chat behavior. The header
carries an optional block (additive evolution, no feature bit):
```json
"tokenizer_config": {
"chat_template": "<Jinja template from chat_template.jinja or tokenizer_config.json>",
"eos_token_ids": [248044, 248045],
"bos_token_id": null,
"pad_token_id": 248055
}
```
The runtime renders the template with HF semantics (trim_blocks,
lstrip_blocks, loop controls, Python string methods) and stops
generation on any id in `eos_token_ids`. Gate:
`tests/chat_template_parity.sh` — the runtime render equals reference
jinja2 byte-for-byte. Files without the block get a ChatML fallback.
## 7. Sparse index
A precomputed bridge "mask → computation skip": active FFN quant groups
(32 neurons each) and heads, per (task, layer) pair.
> Honest status: the engine takes active indices directly from the mask
> bitfields; the index is read and displayed by the CLI but has never
> been used in execution. **Deprecation-pending**: writers SHOULD stop
> emitting it (readers keep parsing existing files); it is revived only
> if the "masks × quantized mmap" path materializes with a measured win.
```
[0 : 4] n_entries : u32
[4 : 8] reserved : u32 (0)
entry (4-aligned):
task_id : u32
layer_idx : u32
n_groups : u32
n_heads : u32
[u16 × n_groups] active FFN-group indices (sorted)
[u8 × n_heads] active head indices (sorted)
zero padding to a multiple of 4
```
A group is active if it contains at least one active mask bit.
## 8. `hash64`
A non-cryptographic 64-bit hash of tensor bytes: murmur3 `fmix64` over
64-bit LE words with positional salt `i·0x9E3779B97F4A7C15`, XOR fold,
`xor len`, final `fmix64`. Bit-for-bit compatible with
`vmfcore.hash64` (Python) and `vmfcore::hash64` (Rust) — hashes of
shared tensors match between `.cmf` and `.vmfc` (backbone dedup across
skill files is free).
Uses: `cortiq verify` (corruption detection), dedup, cache keys.
### 8.1 Section hashes
Metadata integrity (not just tensors):
- Envelope reserve `[0x70:0x78]` = hash64(header JSON), `[0x78:0x80]` =
hash64(directory). Zero = "absent" (older files pass).
- The header JSON carries `section_hashes` — hex hash64 of
masks/vocab/index (u64 as a JSON number would lose precision past
2^53). The header hash in the envelope transitively covers them.
- The envelope itself (first 0x70 bytes) is not hashed: a hash cannot
protect itself; corrupted offsets are caught by bounds/hashes further
down the chain.
- `cortiq verify` checks the whole chain; a single flipped header byte
is an error.
### 8.2 Detached signature (authenticity, opt-in)
The hash chain proves integrity, not authorship. `cortiq sign` writes a
detached `<model>.sig` — JSON `{alg: "ed25519-sha256", pubkey, sha256,
sig}`, Ed25519 over the file's SHA-256 — so the container itself is
never rewritten and old tooling is untouched. `cortiq verify` checks
the signature automatically when the `.sig` sits next to the model;
absence is not an error. Key = a 32-byte hex seed file the signer
keeps private.
## Anti-features — what the format deliberately does NOT have
- **A computable weight layout** — bug class #1 of v1 (writer and reader
"computed" the layout independently and diverged).
- **Silent fallbacks** — v1 would interpret any garbage file as "a 27B
model"; v2 must fail.
- **JSON for bit data** — v1 masks in JSON bloated 3–4×.
- **Declaration fields** — `quality_score: 1.0` by default, area-law
"capacities", Born multipliers in dynamics: a metaphor does not become
a format field until it is measured.
## 9. Skills — a swarm in one file (Patent 15, claims 2/12/15)
One shared backbone + K per-skill records; no record stores a full
model. Storage scales as |backbone| + Σ|deltas|.
**Replacement tensors** are ordinary directory entries named
`skill.{skill_id}.{name_of_replaced_tensor}`, e.g.
`skill.sql.model.layers.3.mlp.gate_proj.weight`. The full logical shape
of the replaced tensor (full-shape — NOT low-rank, NOT a diff list, NOT
a mask), in any encoding of §3. The per-skill delta index (claim 2) is
materialized by the directory: a prefix filter yields skill →
byte-offsets; lazy paging = mmap access to exactly those offsets
(claim 12).
**Registry** — header JSON, additive:
```json
"skills": [{
"id": "sql",
"name": "SQL assistant",
"layers": [3, 4, 5],
"selection": {"metric": "mse", "phi_layer": 20,
"mean": "<f16 base64>", "basis": "<f16 base64>"},
"input_mask_task": null,
"quality": {"metric": "ppl", "backbone": 21.4, "overlaid": 17.9,
"dataset_sha256": "…"}
}]
```
`selection` holds the affine-subspace parameters for recon-argmin
routing (`E = ‖r − BBᵀr‖²/‖φ‖²`, choose the skill with minimal E); the
file is self-sufficient for selection. `quality` is the honest claim-16
contract (overlaid vs backbone on held-out data).
**Execution semantics (claims 1/3/18)**: tensor-source indirection — for
every tensor the runtime reads EITHER the backbone entry OR
`skill.{active}.{name}` if present; replacement instead of addition, a
full per-skill model is never assembled (all tensors are pointers into
one mmap). Soft superposition (claim 14): blended working tensors
`Σwᵢ·Tᵢ`, `wᵢ = softmax(−E/T)`.
**Append-only growth (claim 11)**: adding a skill = appending new
tensors at the file tail + re-emitting directory/header/index at the
tail + updating envelope offsets in place (offset 0 is fixed). Bytes and
offsets of previously written tensors never change; old dir/header bytes
become dead section tails (compatible: readers navigate only through the
envelope). Compaction (`converter/cmf_compact.py`) = a plain rewrite.
### 9.1 Standalone skill files (`SKILL_FILE`, bit 6)
A skill can also travel WITHOUT its backbone: a `.cmf` whose tensor set
is only what a bake changed (plus the mask catalog), bound to the base
it was cut against by identity keys in the registry record:
```json
"skills": [{
"id": "gfx-html",
"layers": [0, 1, "...", 21],
"base_dir_hash": "9f22593eb458bc6f",
"base_arch": "nanbeige",
"task": "specialist",
"provenance": {"corpus": "…", "tensors": 30}
}]
```
- `base_dir_hash` — hex `hash64` of the BASE file's tensor-directory
bytes (the same value the envelope carries at `[0x78]`). A skill is a
delta against exact bytes, not against an architecture: `apply` MUST
refuse a base whose directory hash differs (an explicit `--force`
may override; the result is out of spec).
- `base_arch`, `task`, `provenance` — informative keys: the human check,
the mask-catalog task the skill activates, and where it came from.
Any record with `base_dir_hash` present raises feature bit 6, so a
pre-bit reader refuses the file loudly and a runtime that knows the bit
refuses to RUN it (a partial tensor set is not a model) and points to
`cortiq skill apply <base> <skill> -o out.cmf`, which verifies the key,
overlays tensors and masks over the base, and writes a complete file —
byte-equivalent to the specialist the skill was cut from.
Lifecycle: `skill bake` (specialist) → `skill export --base` (delta +
keys) → publish the small file → `skill apply` on any copy of the base.
Status: fully implemented and gated (container + indirection,
production recipes, recon-argmin routing, append-only + compaction,
soft-blend); claim 16 met by measurement (−24.9% task-PPL in the
runtime).
## 10. Sharding — a model in N files
Naming: `{base}-{no:05}-of-{count:05}.cmf` (spiritually compatible with
safetensors). The user opens ANY name; the runtime normalizes to shard 1
and picks up siblings by pattern.
**Every shard is a standalone valid .cmf**: full envelope, header JSON,
a directory of ITS OWN tensors, its own data blob, its own hashes
(`section_hashes` + per-tensor). `cortiq verify` works on any single
shard without its siblings.
Each shard's header carries:
```json
"shard": { "no": 1, "count": 5 }
```
No block = an ordinary single file (backward compatible: old readers see
shard 1 as a valid but incomplete model and fail honestly on the missing
tensor).
**Content distribution**: tensors are split greedily in canonical order
(`--shard-max-gb` threshold, rough f32 size); the masks/vocab/sparse
index sections, `tokenizer_config` (chat bundle) and the `skills`
registry live ONLY in shard 1 — the rest have empty sections and
`tokenizer_config: null`. Skill tensors (`skill.{id}.*`) are distributed
as ordinary directory entries — the shard-1 registry references them by
name through the merged directory.
**Loading** (`CmfModel::open_sharded`): open shard 1 → mmap all siblings
→ merge directories (each entry remembers its shard index — a runtime
field, never written to disk) → the runtime then works as with a single
file. Errors: opening a non-first shard directly, a missing sibling, a
`count` mismatch.
Gate (Qwen3.5-0.8B q8_2f, 5 shards ≤ 0.6 GB): sharded PPL == unsharded
byte-exactly on the same binary; `verify` green on every shard alone.
## 11. Defragmentation — physical pruning (USPTO App. 19/452,464, claims 9/10 — [PATENTS.md](../PATENTS.md))
A mask (§5) is **virtual sparsity**: pruned neurons are flagged but still
stored in full (all tasks share one backbone — you cannot physically cut
it until you commit to ONE task). Defragmentation turns virtual sparsity
into **physical compression**: pruned FFN neurons are dropped from the
file — they are **neither stored nor computed**. This is Factory-Hard →
defrag from the DTG-MA application (19/452,464): "bake one mask into the weights" and emit a
standalone compact `.cmf`.
**Representation — no new feature bit, backward compatible.** Physical
pruning is expressed ONLY by smaller tensor shapes in the directory (§3
"no computable layout"; the directory is the sole shape authority). The
runtime derives the FFN size from the tensor shape (`gate_proj.rows()`),
not from `arch.intermediate_size`, so a defragged file is an ordinary
smaller dense model that existing readers load unchanged. The masks
section (§5) is **absent** in a defragged file (the mask is the identity
after pruning). `arch.intermediate_size` becomes nominal (= the per-layer
max); the true size lives in each tensor.
**Per-layer variance — better than the patent.** Because the directory
carries an arbitrary shape per tensor, each layer shrinks to its OWN
live-neuron count. The patent must truncate every layer to `max(active)`
(one bottleneck layer caps the ratio at 80.2% vs. 94% achievable) — CMF
has no such limit.
**Invariants (mandatory):**
- per-layer triple: `gate_proj.rows() == up_proj.rows() ==
down_proj.cols() == inter'ₗ`, and `down_proj.rows() == hidden_size`.
One keep-set indexes all three (gate row i, up row i, down col i are
the same neuron);
- neuron axis: rows (axis 0) for `gate_proj`/`up_proj`, columns (axis 1)
for `down_proj`;
- quant group of 32: the `down_proj` neuron axis is its COLUMNS, and
`vbit`/`q4_block` require `in % 32 == 0`. A `down_proj` whose `inter'`
is not a multiple of 32 is written as `q8_2f` (per-row scale — no
column constraint; the converter downgrades automatically). `gate/up`
drop rows, so their columns (= hidden) are unaffected;
- NOT a byte truncation: quant scales are per-group/per-row, so pruning
is dequant → gather live neurons → **requant** at the smaller shape
(the `q8_2f` col-field / `vbit` scales of `down_proj` regenerate for
the shrunk column set); tensor hashes are recomputed;
- `hidden_size`, `embed_tokens`, `lm_head`, and norms are untouched
(skill-selection subspaces depend on hidden).
**One task, standalone file.** Defrag is destructive: one `.cmf` bakes
exactly one task. Multi-task serving stays on masks (§5) or per-skill
replacement tensors (§9).
**Provenance (honest contract).** Header `provenance.defrag`:
```jsonc
"defrag": {
"source_skill": "…/skill_ru",
"pre_intermediate": 3072,
"post_intermediate_max": 640,
"kept_per_layer": [608, 640, 512, ...],
"pruned_ratio": 0.803
}
```
Numerically the dense output of a defragged model is IDENTICAL to the
masked output before quantization (a dead neuron contributes `act·0`
under a mask and is simply absent after defrag); after quantization the
only difference comes from quantizing the smaller matrices.
**Scope:** dense FFN neurons here; MoE experts in §11.1. Attention-head
pruning (the head count is a global runtime scalar) is out of scope.
### 11.1 MoE expert defrag (`cortiq moe-defrag`)
The MoE twin of §11, driven by the routing B-field instead of a neuron
mask: expert usage is strongly task-conditional (measured on a 34.7B
coder: the top-64 expert sets for code vs prose overlap with Jaccard
0.25), so a one-task file can drop the experts that task never routes
to. From a `CMF_MOE_STATS` dump (per-layer expert-selection counts over
a task-representative run), keep per layer the smallest top expert set
reaching `--cover` of the recorded routing mass; drop the rest.
**Representation — same philosophy as §11, no feature bit.**
- Kept experts are renumbered into a CONTIGUOUS per-layer prefix
`mlp.experts.0 … mlp.experts.{k−1}` preserving relative order; a
reader enumerates a layer's experts by tensor PRESENCE up to
`arch.moe.num_experts`, which becomes nominal (= the original
count) — mirroring §11's rule for `intermediate_size`.
- The router tensor's rows are gathered to match, in the same order:
`mlp.gate.weight` becomes `[kept_l, hidden]`, and
`router.rows() == (number of expert entries present)` is a load-time
invariant. `top_k` clamps to the per-layer expert count.
- Selection semantics are unchanged (§2.2): the softmax simply
renormalizes over the kept set. The identical restriction can be
applied at RUNTIME without rewriting the file
(`CMF_MOE_MASK=<stats.json>` + `CMF_MOE_MASK_COVER`) — the two are
mathematically equal, which is how a cover level is perplexity-gated
before committing to the cut.
- Expert payloads are copied verbatim (no requant — the expert axis is
whole tensors, not quant groups), so the surviving weights are
byte-identical to the source and the rewrite streams from the source
mmap.
Measured reference (KAT-Coder 34.7B-A3B, code-calibrated, cover 0.95):
19.6 → 12.7 GB (−35%), held-out code perplexity +2.8%, and on a 24 GB
machine — where the full model paged — decode ×1.8, prefill ×3.3.
Off-task quality degrades by design; like §11, one defragged file bakes
one task. Multi-task serving stays on the full expert set.
**Producing it (native Rust):**
```
cortiq convert --model <hf_dir_or_repo> --defrag <skill_dir> \
--quant q8_2f --output model.cmf
```
`<skill_dir>` carries baked FFN overlays (`tensors/*.npy`) and, if
available, a keep-set `ffn_keep.npy` (bool `[n_layers, intermediate]`,
True = live) from the pruning pipeline. Without `ffn_keep.npy` the
keep-set is autodetected from all-zero `down_proj` columns (the
Factory-Hard bake). The mask-training / bake step lives in the private
research pipeline; the public tool only consumes its artifacts.
## 12. Pipeline containers — text-to-image in one file
The same envelope/directory/blob machinery carries non-LLM pipelines.
The only differences are the `arch_name` tag and namespaced tensor
names; no new sections, no feature bit (a reader that does not execute
the pipeline still validates and inspects the file).
Current instance — `arch_name: "lumina2-image"` (Lumina-Image 2.0,
`cortiq imagine-pack` / `cortiq imagine`): one file packs the whole
text-to-image stack.
- **Namespaces**: `te.*` — the text-encoder transformer (a Gemma-2
class LLM; the header's `arch` block describes THIS component, so
generic tooling reads meaningful dimensions), `dit.*` — the Next-DiT
denoiser, `vae.*` — the VAE decoder. Component config JSONs ride as
`{prefix}.config_json` u8 tensors — the file is self-sufficient.
- **Quantization**: per-tensor as always (§3 directory is the truth) —
typically q4t/q8 matrices for te/dit, f16 for VAE convolutions and
norms.
- **Tokenizer section** (§6) carries the text encoder's tokenizer;
`provenance.pipeline` + `provenance.components` name the recipe.
The measured reference lives in the README (512 px on CPU in minutes,
Metal whole-DiT-block graph on Apple silicon; the wgpu path serves
discrete cards and phones).
---
*Related: [COMPARISON.md](COMPARISON.md) (CMF vs. other model formats),
[project README](../README.md) (overview and quick start),
`python/cmf_reader.py` (standalone reader: stdlib + numpy, reads every
dtype, shards, skills, verify).*
|