File size: 37,682 Bytes
7d1676b
 
e9433fd
7d1676b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1a0c6c0
 
7d1676b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1a0c6c0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7d1676b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1a0c6c0
7d1676b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1a0c6c0
 
7d1676b
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
# CMF v2 — Format Specification

*Languages: **English** · [Русский](SPEC.ru.md) · [中文](SPEC.zh.md)*

**Cortiq Model Format** — a single file carrying everything needed for
sparse, task-routed inference: quantized weights, tokenizer, per-task
masks, a precomputed sparse index — and, uniquely, a **swarm of skills**
sharing one backbone (Patent 15).

> Normative source: this document. Reference
> implementations: Rust reader/runtime (`crates/cortiq-core`,
> `crates/cortiq-engine`), Python writer (`converter/`), and a
> standalone Python reader (`python/cmf_reader.py`, stdlib + numpy).

Three requirements, in priority order:

1. **Correct.** No silent corruption modes: strict magic, version,
   `required_features`, bounds on every section, a 64-bit hash for every
   tensor. A file is either valid or open() returns an error — there is
   no third state.
2. **Fast.** The weight section is page-aligned for mmap, every tensor
   is 64-byte aligned (zero-copy SIMD), the tensor directory is binary —
   read without parsing. Cold (masked-out) weights cost no RSS.
3. **Compact.** Masks are bit-packed (1 bit per neuron), weights are
   q4/q8/variable-bit, the whole file is addressed by one 128-byte
   envelope.

Harmony comes not from feature count but from **a single canon**: one
layout per level (envelope, directory, quant block, mask), byte-for-byte
compatible with the validated `.vmfc` v2 format where the domains
overlap (tensor directory, quant layouts, `hash64`). Never two
definitions of the same thing.

Physical basis (VMF): the model is a vacuum condensate 𝒲; a skill is its
regular core above a critical density; a task mask selects an active
subset without changing weights. The format carries the consequences of
that physics (two-field 𝒲×θ quantization, Born importance, critical mask
threshold) — but **only those confirmed by measurement**.

---

## 1. Envelope (fixed 128 bytes)

All integers are little-endian.

```
[0x00 : 0x04]  magic              = b"CMF\x01"  (4 bytes)
[0x04 : 0x08]  version            : u32 = 2
[0x08 : 0x0C]  flags              : u32 (reserved, 0)
[0x0C : 0x10]  required_features  : u32 (bitmask, §1.1)
[0x10 : 0x18]  header_off         : u64 (= 128)
[0x18 : 0x20]  header_len         : u64   — JSON header (§2)
[0x20 : 0x28]  dir_off            : u64   — tensor directory (§3)
[0x28 : 0x30]  dir_len            : u64
[0x30 : 0x38]  data_off           : u64   — weight blob; multiple of 4096 (§4)
[0x38 : 0x40]  data_len           : u64
[0x40 : 0x48]  masks_off          : u64   — masks section (§5); 0 = absent
[0x48 : 0x50]  masks_len          : u64
[0x50 : 0x58]  vocab_off          : u64   — tokenizer (§6); 0 = absent
[0x58 : 0x60]  vocab_len          : u64
[0x60 : 0x68]  index_off          : u64   — sparse index (§7); 0 = absent
[0x68 : 0x70]  index_len          : u64
[0x70 : 0x80]  reserved           : 16 bytes (§8.1: header/dir hashes)
```

Section order on disk: envelope → header JSON → directory → **weight
blob (aligned to 4096)** → masks → vocab → sparse index. A reader MUST
address sections ONLY through the envelope, never by assuming order.

### 1.1 `required_features`

A bit the reader does not know → `UnsupportedFeature` error (fail-fast;
no "read as best we can").

| bit | name           | meaning |
|-----|----------------|---------|
| 0   | `TENSOR_DIR`   | binary tensor directory (always set in v2) |
| 1   | `BINARY_MASKS` | masks section (§5) present |
| 2   | `QUANT_2F`     | directory contains `q8_2f`/`vbit` tensors (two-field 𝒲×θ quant) |
| 3   | `DELTA_MASKS`  | reserved: XOR mask deltas from a parent |
| 4   | `HOT_PACKS`    | reserved: materialized dense slices |
| 5   | `LOOP_MASKS`   | mask rows are per VISIT (physical layers × loops, pass-major) — a Looped Transformer's two passes carry independent masks (§5.1) |
| 6   | `SKILL_FILE`   | the file is a STANDALONE SKILL: a partial tensor set cut against a specific base, bound by `SkillRecord.base_dir_hash` (§9.1). Not runnable — attach with `cortiq skill apply` |

Unknown **header-JSON** fields are ignored (additive evolution);
breaking changes go only through feature bits or a `version` bump.

### 1.2 Validation rules (normative)

The reader MUST return an error (not a default, not a warning) when:

- magic ≠ `CMF\x01``InvalidMagic`;
- `version` ≠ 2 → `UnsupportedVersion` (v1 is dead: no real v1 files
  exist, no support program will be started);
- an unknown `required_features` bit is set → `UnsupportedFeature`;
- any section extends past EOF, `data_off` is not a multiple of 4096, a
  tensor's `off + nbytes` exceeds `data_len``Bounds`;
- a tensor name is not UTF-8, dtype is unknown, `ndim > 6``Parse`.

Tensor-hash verification is on demand (`cortiq verify`, a loader flag),
not on every open: mmap pages are read lazily.

## 2. Header JSON

UTF-8 JSON, unaligned. Machine-critical data lives in binary sections;
JSON carries architecture and provenance — the parts a human reads.

```jsonc
{
  "format": "cmf",
  "version": 2,
  "arch": {
    "arch_name": "qwen3.5",
    "hidden_size": 5120, "intermediate_size": 17408,
    "num_layers": 64, "num_attention_heads": 24, "num_kv_heads": 4,
    "head_dim": 256, "vocab_size": 248320,
    "layer_types": ["LinearAttention", "...", "FullAttention"],
    "rms_norm_eps": 1e-6,
    "norm_style": "qwen",            // "qwen": x̂·w | "gemma": x̂·(1+w)
    "rope_theta": 1000000.0,
    "yarn": {                         // optional global YaRN profile
      "factor": 128.0, "original_max_position_embeddings": 8192,
      "beta_fast": 32.0, "beta_slow": 1.0, "attention_factor": 1.485203
    },
    "attention_heads_per_layer": [48, 72, 72, 72], // optional; length = num_layers
    "sliding_window": 512,
    "rope_local_base_freq": 10000.0,
    "local_partial_rotary_factor": 1.0,
    "tie_word_embeddings": false,
    "max_position_embeddings": 262144,
    "linear_conv_kernel_dim": 4,
    "linear_num_key_heads": 16, "linear_num_value_heads": 48
  },
  "quant_type": "Q4_BLOCK",          // informational default; truth = per-tensor dtype in the directory
  "provenance": { "tool": "…", "source_model": "…" }   // optional, free-form
}
```

`norm_style` is mandatory for an engine: Gemma-style `(1+w)` applied to
Qwen weights is silent garbage across all ~130 normalizations of a
forward pass.

Capability dispatch is **tensor-presence driven**: an engine decides
per-layer operators by what exists in the directory (q/k biases,
qk-norms, output gate by projection width, MoE router, GDN projections)
— not by matching model names. New models of a known family load with
zero engine changes.

An explicit `SlidingAttention` layer tag selects causal windowed GQA even
when the local/global schedule is irregular. Such layers use
`sliding_window`, `rope_local_base_freq`, and
`local_partial_rotary_factor`; global `FullAttention` layers use
`rope_theta`, `partial_rotary_factor`, and optional `yarn`. The optional
`attention_heads_per_layer` array overrides the base Q-head count for each
layer. Attention projection gating is tensor-presence driven:
`self_attn.g_proj.weight [num_heads, hidden]` means per-head
`softplus(g_proj·x)` gating immediately before `o_proj` (a
`[num_heads·head_dim, hidden]` projection means per-channel gating).
These fields and tensor semantics cover Laguna without introducing a
model-name-specific execution operator.

### 2.1 MTP — multi-token prediction (optional)

If the model carries an MTP head (DeepSeek/Qwen style), arch declares:

```jsonc
"mtp": { "num_layers": 1, "share_lm_head": true, "share_embed": true }
```

MTP tensors are ordinary directory entries under canonical names
(`model.mtp.*`): `enorm.weight`, `hnorm.weight`,
`eh_proj.weight [hidden, 2·hidden]`, `layers.{i}.*` (a standard
transformer block), `norm.weight`.

Semantics: `x = eh_proj·[enorm(embed(t_{p+1})); hnorm(h_p)]` — embedding
FIRST (oracle-verified: the reverse order yields exactly 0% acceptance)
→ block → shared lm_head → draft of token `t_{p+2}`. A reader is not
required to execute MTP (metadata + ordinary tensors, additive
evolution, no feature bit); the CMF runtime uses the head for
speculative decode with a strict guarantee: **output is exactly equal to
plain greedy** — a rejected draft is rolled back from KV.

### 2.2 MoE — mixture-of-experts FFN (optional)

If the model carries MoE layers (Qwen2-MoE / Qwen3-MoE / Qwen3.5-MoE),
arch declares:

```jsonc
"moe": {
  "num_experts": 256, "top_k": 8, "moe_intermediate_size": 512,
  "norm_topk_prob": true,                       // Qwen2-MoE: false
  "shared_expert_intermediate_size": 512,       // absent if no shared expert
  "router_sigmoid": true,                       // optional; default = softmax
  "routed_scaling_factor": 2.5                  // optional; default = 1
}
```

Tensors are ordinary directory entries under HF names:

```
model.layers.{i}.mlp.gate.weight                    [num_experts, hidden]  router
model.layers.{i}.mlp.experts.{e}.{gate,up,down}_proj.weight
model.layers.{i}.mlp.expert_bias                    [num_experts] selection only
model.layers.{i}.mlp.shared_expert.{gate,up,down}_proj.weight
model.layers.{i}.mlp.shared_expert_gate.weight      [1, hidden] optional
```

Which layers are MoE is decided by the PRESENCE of the router in the
directory (per-layer, not per-model): Qwen2-MoE's
`mlp_only_layers`/`decoder_sparse_step` produce mixed models, and dense
layers keep ordinary `mlp.*_proj`.

Execution semantics (HF parity, gated by `tests/moe_parity.sh` across
multiple families): by default, softmax over ALL router logits; when
`router_sigmoid`, score each expert independently with sigmoid. An optional
`expert_bias` affects top-k selection only, not the gathered weights. Select
top-k (ties: lower index), optionally renormalize the selected weights, then
apply `routed_scaling_factor` and compute Σwₑ·FFNₑ(x). The shared expert is
always added: with weight `sigmoid(shared_expert_gate·x)` when that tensor is
present, otherwise with weight 1 (Laguna).
Experts stay quantized in mmap; per token only the pages of the selected
k are touched — the same residency story as skills. Writers SHOULD lay
a layer's expert tensors out role-contiguously (all `gate_proj` of
experts 0…N−1 back to back, then all `up_proj`, then all `down_proj`)
— GPU backends can then treat a layer's expert bank as one region
instead of gathering hundreds of slices; the native importer and
`moe-defrag` both emit this order. Each expert is a
separate directory entry with ITS OWN dtype: that is the carrier of
per-expert bit allocation (P15 claim 12) — implemented, gated by
`tests/moe_vbit.sh`; the B-field (router selection frequencies via
`--route-stats`) was measured end-to-end on a 35B model.

## 3. Tensor directory

Byte-for-byte the `.vmfc` v2 layout (single canon, shared reference
parser):

```
[0 : 8 ]  count    : u64
[8 : 16]  pool_off : u64            (name-pool offset from section start)
[16 : 16 + count·56]  56-byte records:
   name_off : u32   (relative to pool_off)
   name_len : u16
   dtype    : u8    (§3.1)
   ndim     : u8    (≤ 6)
   shape    : u32 × 6  (zero-padded tail)
   off      : u64   (RELATIVE to data_off; multiple of 64)
   nbytes   : u64
   hash     : u64   (hash64 of the tensor bytes, §8)
[pool_off : …]  UTF-8 name pool
```

Tensor names are **1:1 with the source model**
(`model.layers.{i}.mlp.gate_proj.weight`, `model.embed_tokens.weight`,
`lm_head.weight`, …). The format does not prescribe a tensor set: the
directory is the single source of truth for what the blob contains.
There is no "computable layout".

### 3.1 `dtype`

Numbering shared with `.vmfc` (ids are never reused):

| id | name       | status in CMF v2 |
|----|-----------|------------------|
| 0  | `f32`     | ✅ read/write |
| 1  | `f16`     | ✅ read/write (norms and 1-D are always f16) |
| 2  | `bf16`    | ✅ read/write |
| 3  | `q8_row`  | ✅ read/write |
| 4  | `q4_block`| ✅ read/write |
| 5  | `mix8_4`  | reserved |
| 6  | `u8`      | reserved |
| 7  | `q4_col`  | reserved |
| 8  | `vbit`    | ✅ read/write (`QUANT_2F` bit), variable 3–8 bit |
| 9  | `q8_2f`   | ✅ read/write (`QUANT_2F` bit), 𝒲×θ |
| 10 | `vbit_ro` | ✅ read/write — `vbit` + in-file row-offset table (O(1) row access); converter default for `--quant vbit` |
| 11 | `q4_tiled`| ✅ read/write — q4 in interleaved `[f16 scale][16B nibbles]` tiles (`--quant q4t`) |
| 12 | `q1`      | ✅ read/write — 1-bit binary, for 1-bit-TRAINED models only (`--quant q1`) |
| 13 | `q1s`     | ✅ read/write — `q1` base + sparse high-precision outlier overlay (1-bit PTQ of normal checkpoints) |
| 14 | `q1t`     | ✅ read/write — ternary `{−s, 0, +s}` base-3 tiles + per-row outlier overlay (~2.25 bpw + overlay) |
| 15 | `q4tp`    | ✅ read/write — `q4_tiled` nibbles with the per-tile scale as a 5-bit rung on a per-row ladder (`--quant q4tp`, or `requant` in place) |

### 3.2 Quant layouts (canon = `.vmfc`: "quants first, then scales")

- **`q8_row`** (2-D `[out, in]` only):
  `[int8 : out·in][f16 : out]` — one scale per row,
  `w = q[o,i]·scale[o]`, `scale[o] = absmax(row_o)/127`.
- **`q4_block`**: groups of 32 over the flattened tensor, zero-padded;
  `[u8 : ceil(n/32)·16][f16 : ceil(n/32)]`.
  Nibbles: element `2k` low, `2k+1` high; `w = (q − 8)·scale`,
  `scale = absmax(group)/7`.
- **1-D tensors and tensors < 32 elements are always `f16`**
  (normalization precision at maximal matrix compression).
- **`q8_2f`**: `[int8][f16 row-scale][f16 col-field]`,
  `w = q·scale[o]·col[i]` — the two-field Madelung split 𝒲×θ, validated
  in vmfcore (+37% at equal size; recovers ~75% of the q8→f16 gap on
  outlier input channels).
- **`vbit`** (2-D only, `in % 32 == 0`; P13 FIG.3):
  `[u8 bits: rows][f16 scales: rows·in/32][bit-packed rows, MSB-first,
  each row padded to a byte]`; `w = (u − L)·scale[r,g]`,
  `L = 2^{b−1}−1`, levels b ∈ {3,4,5,6,8}, floor 3 (claim 13).
  Allocation b_r: water-filling over the log2 row amplitude toward the
  tensor's mean budget; for MoE experts the budget is SHARED across the
  family (layer × projection): the shift `ā_expert − ā_family` is
  equivalent to joint water-filling over all experts' rows — a loud
  expert gets more bits, a quiet one is pinned to the floor (P15
  claim 12; gate `tests/moe_vbit.sh`). Optionally the allocation takes
  the product with a B-field — router selection frequencies collected
  at calibration (`b ∝ log2(A·B)`, truncated Fisher).
- **`vbit_ro`** (2-D only, `in % 32 == 0`): the same bits/scales/packed
  encoding as `vbit`, plus `u32 row_offsets[rows+1]` (relative to the
  packed area) between the scales and the packed rows —
  `[u8 bits: rows][f16 scales: rows·in/32][u32 offsets: rows+1][packed]`.
  Readers get O(1) row access without a prefix scan over bit widths.
  The byte semantics of `vbit = 8` are untouched; new id on purpose.
- **`q4_tiled`** (2-D only, `in % 32 == 0`):
  `repeat per 32-group { [f16 scale][16B nibbles] }` — 18-byte tiles,
  one sequential memory stream instead of two distant ones. Values and
  nibble order are identical to `q4_block`; only the placement of the
  scale differs (kernel-measured ×1.66 ARM / ×1.13 AVX2 over split).
- **`q4tp`** (2-D only, `in % 32 == 0`):
  `[nibbles: rows·gpr·16][row params: rows × (f16 lo, f16 step)]
   [codes: rows × ceil(gpr·5/8), 5-bit LSB-first, row-aligned]`,
  `gpr = in/32`. A tile's scale is `2^(lo[r] + code·step[r])`, so a reader
  expands one row's 32-rung ladder once and then reads scales by table
  lookup. Nibble values and order are identical to `q4_tiled`; only the
  scale's representation differs. 4.17 bits/weight against 4.50 — the
  f16 scale was 11% of a q4t file.
  `lo`/`step` come from the row's exact min/max log-scale, so no code is
  ever out of range and the format needs no escape hatch. Encoders MUST
  round `lo`/`step` to f16 **before** choosing codes, and MUST quantize the
  nibbles against the reconstructed scale — otherwise writer and reader
  disagree, the same trap that makes a q4 encoder round its scale first.
- **`q1`** (2-D only, `in % 32 == 0`):
  `repeat per 32-group { [f16 scale][4B sign bits] }` — 6-byte tiles,
  1.5 bits/weight. Bit k of byte j (LSB-first) is weight j·8+k of the
  group; `w = scale·(2·bit − 1) ∈ {−s, +s}`, `scale = mean|group|`
  (the L2-optimal binary level). Intended for 1-bit-TRAINED models
  (Bonsai / BitNet class), where per-group weights already sit on two
  levels and the encoding is lossless up to f16; as post-training
  quantization of a normal checkpoint it destroys quality, so
  converters expose it only as an explicit opt-in.
- **`q1s`** (2-D only, `in % 32 == 0`): a `q1` base (identical 6-byte
  tiles; outliers are EXCLUDED from the group scale) followed by a
  sparse high-precision overlay: `[u32 count]` then
  `count × { [u32 flat-index][f16 value] }` — the salient weights kept
  at full precision (holographic transfer / SpQR-style) and restored
  verbatim at dequant. Variable length: `expected_nbytes` is
  undefined, the reader trusts the directory's stored span. Lets a
  NORMAL checkpoint survive 1-bit where plain `q1` cannot.
- **`q1t`** (2-D only, `in % 32 == 0`, `in` must fit `u16`): ternary
  BitNet-b1.58-style `{−s, 0, +s}`. Base:
  `repeat per 32-group { [f16 scale][7B base-3 codes] }` — 9-byte
  tiles, 5 ternary values per byte (3⁵ = 243 ≤ 256; code 0 → 0,
  1 → +s, 2 → −s), ~2.25 bits/weight. Then a per-row outlier overlay:
  `[u32 row_ptr[rows+1]]` followed by `{ [u16 col][f16 value] }`
  entries grouped by row (row `r`'s outliers are
  `[row_ptr[r], row_ptr[r+1])`; `col` is a within-row index) — 4
  bytes per outlier, no binary search. Capturing the many near-zero
  weights exactly is the decisive PTQ win over binary. Variable
  length, same span rule as `q1s`.

## 4. Weight blob

`data_off` is a multiple of 4096 (page-aligned mmap); every tensor
inside starts on a 64-byte boundary (SIMD loads, cache lines). Zero
padding between tensors. A reader interprets the blob only through the
directory.

## 5. Masks section

A task mask = bit fields of "what is active" over shared weights
(weights do not change — the VMF principle: a skill selects a subset of
the condensate).

```
[0 : 4]  n_masks  : u32
[4 : 8]  meta_len : u32
[8 : 8 + meta_len]  JSON meta (§5.1)
[…]      mask blobs, each aligned to 8 from the section start
```

One mask blob (sizes derived from arch, no internal headers):

```
[n_layers × ffn_bytes]   FFN bitfields      ffn_bytes  = ceil(intermediate_size / 8)
[n_layers × head_bytes]  head bitfields     head_bytes = ceil(num_attention_heads / 8)
[gates_bytes]            layer_gates        gates_bytes = ceil(num_layers / 8)
[n_layers × expert_bytes] expert bitfields  OPTIONAL — only when the mask's meta
                                            sets "has_expert_fields": true;
                                            expert_bytes = ceil(moe.num_experts / 8)
```

Bit order is LSB-first: neuron `i` = bit `i % 8` of byte `i / 8`; bit
set → active. **Tail bits beyond the dimension MUST be zero** (or
popcount sees phantom neurons/heads).

The optional expert area (additive: old readers never look past the
gates, and each mask's `blob_len` is explicit) makes a task mask narrow
MoE ROUTING: bit `e` of layer `l`'s row set → expert `e` is routable
for this task; selection then happens over the routable set only, the
router softmax renormalizing over it. This is the runtime-switchable
twin of §11.1's physical expert defrag — one file with the full expert
set serves many specialists (`cortiq moe-mask` writes such masks,
`run --task <name>` activates one; verified token-identical to the
equivalent runtime restriction). A layer whose row is all-ones is
unrestricted; a mask without the area restricts nothing.

### 5.1 Mask JSON meta

```jsonc
{
  "default_task": "general",
  "masks": [{
    "task_id": 0, "name": "general", "description": null,
    "sparsity": 0.62,
    "quality": {                    // null = NOT MEASURED (declaring 1.0 is forbidden)
      "metric": "heldout_ppl_ratio", "value": 0.97,
      "baseline_dense": 6.10, "n_samples": 512, "dataset_sha256": "…"
    },
    "parent": null, "priority": "Fallback", "has_hot_pack": false,
    "blob_off": 4096, "blob_len": 139328   // relative to section start
  }]
}
```

`quality` is a **held-out contract**, not a declaration: a converter
without a measured metric writes `null`; the runtime logs a warning when
switching to an unmeasured mask.

## 6. Tokenizer section

The bytes of HuggingFace `tokenizer.json`, verbatim. The model is
self-contained: one file = one unit of distribution. A sidecar file
remains a debugging fallback.

### 6.1 Chat bundle (`header.tokenizer_config`)

The file — not the runtime binary — defines chat behavior. The header
carries an optional block (additive evolution, no feature bit):

```json
"tokenizer_config": {
  "chat_template": "<Jinja template from chat_template.jinja or tokenizer_config.json>",
  "eos_token_ids": [248044, 248045],
  "bos_token_id": null,
  "pad_token_id": 248055
}
```

The runtime renders the template with HF semantics (trim_blocks,
lstrip_blocks, loop controls, Python string methods) and stops
generation on any id in `eos_token_ids`. Gate:
`tests/chat_template_parity.sh` — the runtime render equals reference
jinja2 byte-for-byte. Files without the block get a ChatML fallback.

## 7. Sparse index

A precomputed bridge "mask → computation skip": active FFN quant groups
(32 neurons each) and heads, per (task, layer) pair.

> Honest status: the engine takes active indices directly from the mask
> bitfields; the index is read and displayed by the CLI but has never
> been used in execution. **Deprecation-pending**: writers SHOULD stop
> emitting it (readers keep parsing existing files); it is revived only
> if the "masks × quantized mmap" path materializes with a measured win.

```
[0 : 4]  n_entries : u32
[4 : 8]  reserved  : u32 (0)
entry (4-aligned):
   task_id   : u32
   layer_idx : u32
   n_groups  : u32
   n_heads   : u32
   [u16 × n_groups]  active FFN-group indices (sorted)
   [u8  × n_heads]   active head indices (sorted)
   zero padding to a multiple of 4
```

A group is active if it contains at least one active mask bit.

## 8. `hash64`

A non-cryptographic 64-bit hash of tensor bytes: murmur3 `fmix64` over
64-bit LE words with positional salt `i·0x9E3779B97F4A7C15`, XOR fold,
`xor len`, final `fmix64`. Bit-for-bit compatible with
`vmfcore.hash64` (Python) and `vmfcore::hash64` (Rust) — hashes of
shared tensors match between `.cmf` and `.vmfc` (backbone dedup across
skill files is free).

Uses: `cortiq verify` (corruption detection), dedup, cache keys.

### 8.1 Section hashes

Metadata integrity (not just tensors):

- Envelope reserve `[0x70:0x78]` = hash64(header JSON), `[0x78:0x80]` =
  hash64(directory). Zero = "absent" (older files pass).
- The header JSON carries `section_hashes` — hex hash64 of
  masks/vocab/index (u64 as a JSON number would lose precision past
  2^53). The header hash in the envelope transitively covers them.
- The envelope itself (first 0x70 bytes) is not hashed: a hash cannot
  protect itself; corrupted offsets are caught by bounds/hashes further
  down the chain.
- `cortiq verify` checks the whole chain; a single flipped header byte
  is an error.

### 8.2 Detached signature (authenticity, opt-in)

The hash chain proves integrity, not authorship. `cortiq sign` writes a
detached `<model>.sig` — JSON `{alg: "ed25519-sha256", pubkey, sha256,
sig}`, Ed25519 over the file's SHA-256 — so the container itself is
never rewritten and old tooling is untouched. `cortiq verify` checks
the signature automatically when the `.sig` sits next to the model;
absence is not an error. Key = a 32-byte hex seed file the signer
keeps private.

## Anti-features — what the format deliberately does NOT have

- **A computable weight layout** — bug class #1 of v1 (writer and reader
  "computed" the layout independently and diverged).
- **Silent fallbacks** — v1 would interpret any garbage file as "a 27B
  model"; v2 must fail.
- **JSON for bit data** — v1 masks in JSON bloated 3–4×.
- **Declaration fields** — `quality_score: 1.0` by default, area-law
  "capacities", Born multipliers in dynamics: a metaphor does not become
  a format field until it is measured.

## 9. Skills — a swarm in one file (Patent 15, claims 2/12/15)

One shared backbone + K per-skill records; no record stores a full
model. Storage scales as |backbone| + Σ|deltas|.

**Replacement tensors** are ordinary directory entries named
`skill.{skill_id}.{name_of_replaced_tensor}`, e.g.
`skill.sql.model.layers.3.mlp.gate_proj.weight`. The full logical shape
of the replaced tensor (full-shape — NOT low-rank, NOT a diff list, NOT
a mask), in any encoding of §3. The per-skill delta index (claim 2) is
materialized by the directory: a prefix filter yields skill →
byte-offsets; lazy paging = mmap access to exactly those offsets
(claim 12).

**Registry** — header JSON, additive:

```json
"skills": [{
  "id": "sql",
  "name": "SQL assistant",
  "layers": [3, 4, 5],
  "selection": {"metric": "mse", "phi_layer": 20,
                 "mean": "<f16 base64>", "basis": "<f16 base64>"},
  "input_mask_task": null,
  "quality": {"metric": "ppl", "backbone": 21.4, "overlaid": 17.9,
               "dataset_sha256": "…"}
}]
```

`selection` holds the affine-subspace parameters for recon-argmin
routing (`E = ‖r − BBᵀr‖²/‖φ‖²`, choose the skill with minimal E); the
file is self-sufficient for selection. `quality` is the honest claim-16
contract (overlaid vs backbone on held-out data).

**Execution semantics (claims 1/3/18)**: tensor-source indirection — for
every tensor the runtime reads EITHER the backbone entry OR
`skill.{active}.{name}` if present; replacement instead of addition, a
full per-skill model is never assembled (all tensors are pointers into
one mmap). Soft superposition (claim 14): blended working tensors
`Σwᵢ·Tᵢ`, `wᵢ = softmax(−E/T)`.

**Append-only growth (claim 11)**: adding a skill = appending new
tensors at the file tail + re-emitting directory/header/index at the
tail + updating envelope offsets in place (offset 0 is fixed). Bytes and
offsets of previously written tensors never change; old dir/header bytes
become dead section tails (compatible: readers navigate only through the
envelope). Compaction (`converter/cmf_compact.py`) = a plain rewrite.

### 9.1 Standalone skill files (`SKILL_FILE`, bit 6)

A skill can also travel WITHOUT its backbone: a `.cmf` whose tensor set
is only what a bake changed (plus the mask catalog), bound to the base
it was cut against by identity keys in the registry record:

```json
"skills": [{
  "id": "gfx-html",
  "layers": [0, 1, "...", 21],
  "base_dir_hash": "9f22593eb458bc6f",
  "base_arch": "nanbeige",
  "task": "specialist",
  "provenance": {"corpus": "…", "tensors": 30}
}]
```

- `base_dir_hash` — hex `hash64` of the BASE file's tensor-directory
  bytes (the same value the envelope carries at `[0x78]`). A skill is a
  delta against exact bytes, not against an architecture: `apply` MUST
  refuse a base whose directory hash differs (an explicit `--force`
  may override; the result is out of spec).
- `base_arch`, `task`, `provenance` — informative keys: the human check,
  the mask-catalog task the skill activates, and where it came from.

Any record with `base_dir_hash` present raises feature bit 6, so a
pre-bit reader refuses the file loudly and a runtime that knows the bit
refuses to RUN it (a partial tensor set is not a model) and points to
`cortiq skill apply <base> <skill> -o out.cmf`, which verifies the key,
overlays tensors and masks over the base, and writes a complete file —
byte-equivalent to the specialist the skill was cut from.

Lifecycle: `skill bake` (specialist) → `skill export --base` (delta +
keys) → publish the small file → `skill apply` on any copy of the base.

Status: fully implemented and gated (container + indirection,
production recipes, recon-argmin routing, append-only + compaction,
soft-blend); claim 16 met by measurement (−24.9% task-PPL in the
runtime).

## 10. Sharding — a model in N files

Naming: `{base}-{no:05}-of-{count:05}.cmf` (spiritually compatible with
safetensors). The user opens ANY name; the runtime normalizes to shard 1
and picks up siblings by pattern.

**Every shard is a standalone valid .cmf**: full envelope, header JSON,
a directory of ITS OWN tensors, its own data blob, its own hashes
(`section_hashes` + per-tensor). `cortiq verify` works on any single
shard without its siblings.

Each shard's header carries:

```json
"shard": { "no": 1, "count": 5 }
```

No block = an ordinary single file (backward compatible: old readers see
shard 1 as a valid but incomplete model and fail honestly on the missing
tensor).

**Content distribution**: tensors are split greedily in canonical order
(`--shard-max-gb` threshold, rough f32 size); the masks/vocab/sparse
index sections, `tokenizer_config` (chat bundle) and the `skills`
registry live ONLY in shard 1 — the rest have empty sections and
`tokenizer_config: null`. Skill tensors (`skill.{id}.*`) are distributed
as ordinary directory entries — the shard-1 registry references them by
name through the merged directory.

**Loading** (`CmfModel::open_sharded`): open shard 1 → mmap all siblings
→ merge directories (each entry remembers its shard index — a runtime
field, never written to disk) → the runtime then works as with a single
file. Errors: opening a non-first shard directly, a missing sibling, a
`count` mismatch.

Gate (Qwen3.5-0.8B q8_2f, 5 shards ≤ 0.6 GB): sharded PPL == unsharded
byte-exactly on the same binary; `verify` green on every shard alone.

## 11. Defragmentation — physical pruning (USPTO App. 19/452,464, claims 9/10 — [PATENTS.md](../PATENTS.md))

A mask (§5) is **virtual sparsity**: pruned neurons are flagged but still
stored in full (all tasks share one backbone — you cannot physically cut
it until you commit to ONE task). Defragmentation turns virtual sparsity
into **physical compression**: pruned FFN neurons are dropped from the
file — they are **neither stored nor computed**. This is Factory-Hard →
defrag from the DTG-MA application (19/452,464): "bake one mask into the weights" and emit a
standalone compact `.cmf`.

**Representation — no new feature bit, backward compatible.** Physical
pruning is expressed ONLY by smaller tensor shapes in the directory (§3
"no computable layout"; the directory is the sole shape authority). The
runtime derives the FFN size from the tensor shape (`gate_proj.rows()`),
not from `arch.intermediate_size`, so a defragged file is an ordinary
smaller dense model that existing readers load unchanged. The masks
section (§5) is **absent** in a defragged file (the mask is the identity
after pruning). `arch.intermediate_size` becomes nominal (= the per-layer
max); the true size lives in each tensor.

**Per-layer variance — better than the patent.** Because the directory
carries an arbitrary shape per tensor, each layer shrinks to its OWN
live-neuron count. The patent must truncate every layer to `max(active)`
(one bottleneck layer caps the ratio at 80.2% vs. 94% achievable) — CMF
has no such limit.

**Invariants (mandatory):**

- per-layer triple: `gate_proj.rows() == up_proj.rows() ==
  down_proj.cols() == inter'ₗ`, and `down_proj.rows() == hidden_size`.
  One keep-set indexes all three (gate row i, up row i, down col i are
  the same neuron);
- neuron axis: rows (axis 0) for `gate_proj`/`up_proj`, columns (axis 1)
  for `down_proj`;
- quant group of 32: the `down_proj` neuron axis is its COLUMNS, and
  `vbit`/`q4_block` require `in % 32 == 0`. A `down_proj` whose `inter'`
  is not a multiple of 32 is written as `q8_2f` (per-row scale — no
  column constraint; the converter downgrades automatically). `gate/up`
  drop rows, so their columns (= hidden) are unaffected;
- NOT a byte truncation: quant scales are per-group/per-row, so pruning
  is dequant → gather live neurons → **requant** at the smaller shape
  (the `q8_2f` col-field / `vbit` scales of `down_proj` regenerate for
  the shrunk column set); tensor hashes are recomputed;
- `hidden_size`, `embed_tokens`, `lm_head`, and norms are untouched
  (skill-selection subspaces depend on hidden).

**One task, standalone file.** Defrag is destructive: one `.cmf` bakes
exactly one task. Multi-task serving stays on masks (§5) or per-skill
replacement tensors (§9).

**Provenance (honest contract).** Header `provenance.defrag`:

```jsonc
"defrag": {
  "source_skill": "…/skill_ru",
  "pre_intermediate": 3072,
  "post_intermediate_max": 640,
  "kept_per_layer": [608, 640, 512, ...],
  "pruned_ratio": 0.803
}
```

Numerically the dense output of a defragged model is IDENTICAL to the
masked output before quantization (a dead neuron contributes `act·0`
under a mask and is simply absent after defrag); after quantization the
only difference comes from quantizing the smaller matrices.

**Scope:** dense FFN neurons here; MoE experts in §11.1. Attention-head
pruning (the head count is a global runtime scalar) is out of scope.

### 11.1 MoE expert defrag (`cortiq moe-defrag`)

The MoE twin of §11, driven by the routing B-field instead of a neuron
mask: expert usage is strongly task-conditional (measured on a 34.7B
coder: the top-64 expert sets for code vs prose overlap with Jaccard
0.25), so a one-task file can drop the experts that task never routes
to. From a `CMF_MOE_STATS` dump (per-layer expert-selection counts over
a task-representative run), keep per layer the smallest top expert set
reaching `--cover` of the recorded routing mass; drop the rest.

**Representation — same philosophy as §11, no feature bit.**

- Kept experts are renumbered into a CONTIGUOUS per-layer prefix
  `mlp.experts.0 … mlp.experts.{k−1}` preserving relative order; a
  reader enumerates a layer's experts by tensor PRESENCE up to
  `arch.moe.num_experts`, which becomes nominal (= the original
  count) — mirroring §11's rule for `intermediate_size`.
- The router tensor's rows are gathered to match, in the same order:
  `mlp.gate.weight` becomes `[kept_l, hidden]`, and
  `router.rows() == (number of expert entries present)` is a load-time
  invariant. `top_k` clamps to the per-layer expert count.
- Selection semantics are unchanged (§2.2): the softmax simply
  renormalizes over the kept set. The identical restriction can be
  applied at RUNTIME without rewriting the file
  (`CMF_MOE_MASK=<stats.json>` + `CMF_MOE_MASK_COVER`) — the two are
  mathematically equal, which is how a cover level is perplexity-gated
  before committing to the cut.
- Expert payloads are copied verbatim (no requant — the expert axis is
  whole tensors, not quant groups), so the surviving weights are
  byte-identical to the source and the rewrite streams from the source
  mmap.

Measured reference (KAT-Coder 34.7B-A3B, code-calibrated, cover 0.95):
19.6 → 12.7 GB (−35%), held-out code perplexity +2.8%, and on a 24 GB
machine — where the full model paged — decode ×1.8, prefill ×3.3.

Off-task quality degrades by design; like §11, one defragged file bakes
one task. Multi-task serving stays on the full expert set.

**Producing it (native Rust):**

```
cortiq convert --model <hf_dir_or_repo> --defrag <skill_dir> \
  --quant q8_2f --output model.cmf
```

`<skill_dir>` carries baked FFN overlays (`tensors/*.npy`) and, if
available, a keep-set `ffn_keep.npy` (bool `[n_layers, intermediate]`,
True = live) from the pruning pipeline. Without `ffn_keep.npy` the
keep-set is autodetected from all-zero `down_proj` columns (the
Factory-Hard bake). The mask-training / bake step lives in the private
research pipeline; the public tool only consumes its artifacts.

## 12. Pipeline containers — text-to-image in one file

The same envelope/directory/blob machinery carries non-LLM pipelines.
The only differences are the `arch_name` tag and namespaced tensor
names; no new sections, no feature bit (a reader that does not execute
the pipeline still validates and inspects the file).

Current instance — `arch_name: "lumina2-image"` (Lumina-Image 2.0,
`cortiq imagine-pack` / `cortiq imagine`): one file packs the whole
text-to-image stack.

- **Namespaces**: `te.*` — the text-encoder transformer (a Gemma-2
  class LLM; the header's `arch` block describes THIS component, so
  generic tooling reads meaningful dimensions), `dit.*` — the Next-DiT
  denoiser, `vae.*` — the VAE decoder. Component config JSONs ride as
  `{prefix}.config_json` u8 tensors — the file is self-sufficient.
- **Quantization**: per-tensor as always (§3 directory is the truth) —
  typically q4t/q8 matrices for te/dit, f16 for VAE convolutions and
  norms.
- **Tokenizer section** (§6) carries the text encoder's tokenizer;
  `provenance.pipeline` + `provenance.components` name the recipe.

The measured reference lives in the README (512 px on CPU in minutes,
Metal whole-DiT-block graph on Apple silicon; the wgpu path serves
discrete cards and phones).

---

*Related: [COMPARISON.md](COMPARISON.md) (CMF vs. other model formats),
[project README](../README.md) (overview and quick start),
`python/cmf_reader.py` (standalone reader: stdlib + numpy, reads every
dtype, shards, skills, verify).*