CMF v2 — Format Specification
Languages: English · Русский · 中文
Cortiq Model Format — a single file carrying everything needed for sparse, task-routed inference: quantized weights, tokenizer, per-task masks, a precomputed sparse index — and, uniquely, a swarm of skills sharing one backbone (Patent 15).
Normative source: this document. Reference implementations: Rust reader/runtime (
crates/cortiq-core,crates/cortiq-engine), Python writer (converter/), and a standalone Python reader (python/cmf_reader.py, stdlib + numpy).
Three requirements, in priority order:
- Correct. No silent corruption modes: strict magic, version,
required_features, bounds on every section, a 64-bit hash for every tensor. A file is either valid or open() returns an error — there is no third state. - Fast. The weight section is page-aligned for mmap, every tensor is 64-byte aligned (zero-copy SIMD), the tensor directory is binary — read without parsing. Cold (masked-out) weights cost no RSS.
- Compact. Masks are bit-packed (1 bit per neuron), weights are q4/q8/variable-bit, the whole file is addressed by one 128-byte envelope.
Harmony comes not from feature count but from a single canon: one
layout per level (envelope, directory, quant block, mask), byte-for-byte
compatible with the validated .vmfc v2 format where the domains
overlap (tensor directory, quant layouts, hash64). Never two
definitions of the same thing.
Physical basis (VMF): the model is a vacuum condensate 𝒲; a skill is its regular core above a critical density; a task mask selects an active subset without changing weights. The format carries the consequences of that physics (two-field 𝒲×θ quantization, Born importance, critical mask threshold) — but only those confirmed by measurement.
1. Envelope (fixed 128 bytes)
All integers are little-endian.
[0x00 : 0x04] magic = b"CMF\x01" (4 bytes)
[0x04 : 0x08] version : u32 = 2
[0x08 : 0x0C] flags : u32 (reserved, 0)
[0x0C : 0x10] required_features : u32 (bitmask, §1.1)
[0x10 : 0x18] header_off : u64 (= 128)
[0x18 : 0x20] header_len : u64 — JSON header (§2)
[0x20 : 0x28] dir_off : u64 — tensor directory (§3)
[0x28 : 0x30] dir_len : u64
[0x30 : 0x38] data_off : u64 — weight blob; multiple of 4096 (§4)
[0x38 : 0x40] data_len : u64
[0x40 : 0x48] masks_off : u64 — masks section (§5); 0 = absent
[0x48 : 0x50] masks_len : u64
[0x50 : 0x58] vocab_off : u64 — tokenizer (§6); 0 = absent
[0x58 : 0x60] vocab_len : u64
[0x60 : 0x68] index_off : u64 — sparse index (§7); 0 = absent
[0x68 : 0x70] index_len : u64
[0x70 : 0x80] reserved : 16 bytes (§8.1: header/dir hashes)
Section order on disk: envelope → header JSON → directory → weight blob (aligned to 4096) → masks → vocab → sparse index. A reader MUST address sections ONLY through the envelope, never by assuming order.
1.1 required_features
A bit the reader does not know → UnsupportedFeature error (fail-fast;
no "read as best we can").
| bit | name | meaning |
|---|---|---|
| 0 | TENSOR_DIR |
binary tensor directory (always set in v2) |
| 1 | BINARY_MASKS |
masks section (§5) present |
| 2 | QUANT_2F |
directory contains q8_2f/vbit tensors (two-field 𝒲×θ quant) |
| 3 | DELTA_MASKS |
reserved: XOR mask deltas from a parent |
| 4 | HOT_PACKS |
reserved: materialized dense slices |
| 5 | LOOP_MASKS |
mask rows are per VISIT (physical layers × loops, pass-major) — a Looped Transformer's two passes carry independent masks (§5.1) |
| 6 | SKILL_FILE |
the file is a STANDALONE SKILL: a partial tensor set cut against a specific base, bound by SkillRecord.base_dir_hash (§9.1). Not runnable — attach with cortiq skill apply |
Unknown header-JSON fields are ignored (additive evolution);
breaking changes go only through feature bits or a version bump.
1.2 Validation rules (normative)
The reader MUST return an error (not a default, not a warning) when:
- magic ≠
CMF\x01→InvalidMagic; version≠ 2 →UnsupportedVersion(v1 is dead: no real v1 files exist, no support program will be started);- an unknown
required_featuresbit is set →UnsupportedFeature; - any section extends past EOF,
data_offis not a multiple of 4096, a tensor'soff + nbytesexceedsdata_len→Bounds; - a tensor name is not UTF-8, dtype is unknown,
ndim > 6→Parse.
Tensor-hash verification is on demand (cortiq verify, a loader flag),
not on every open: mmap pages are read lazily.
2. Header JSON
UTF-8 JSON, unaligned. Machine-critical data lives in binary sections; JSON carries architecture and provenance — the parts a human reads.
{
"format": "cmf",
"version": 2,
"arch": {
"arch_name": "qwen3.5",
"hidden_size": 5120, "intermediate_size": 17408,
"num_layers": 64, "num_attention_heads": 24, "num_kv_heads": 4,
"head_dim": 256, "vocab_size": 248320,
"layer_types": ["LinearAttention", "...", "FullAttention"],
"rms_norm_eps": 1e-6,
"norm_style": "qwen", // "qwen": x̂·w | "gemma": x̂·(1+w)
"rope_theta": 1000000.0,
"yarn": { // optional global YaRN profile
"factor": 128.0, "original_max_position_embeddings": 8192,
"beta_fast": 32.0, "beta_slow": 1.0, "attention_factor": 1.485203
},
"attention_heads_per_layer": [48, 72, 72, 72], // optional; length = num_layers
"sliding_window": 512,
"rope_local_base_freq": 10000.0,
"local_partial_rotary_factor": 1.0,
"tie_word_embeddings": false,
"max_position_embeddings": 262144,
"linear_conv_kernel_dim": 4,
"linear_num_key_heads": 16, "linear_num_value_heads": 48
},
"quant_type": "Q4_BLOCK", // informational default; truth = per-tensor dtype in the directory
"provenance": { "tool": "…", "source_model": "…" } // optional, free-form
}
norm_style is mandatory for an engine: Gemma-style (1+w) applied to
Qwen weights is silent garbage across all ~130 normalizations of a
forward pass.
Capability dispatch is tensor-presence driven: an engine decides per-layer operators by what exists in the directory (q/k biases, qk-norms, output gate by projection width, MoE router, GDN projections) — not by matching model names. New models of a known family load with zero engine changes.
An explicit SlidingAttention layer tag selects causal windowed GQA even
when the local/global schedule is irregular. Such layers use
sliding_window, rope_local_base_freq, and
local_partial_rotary_factor; global FullAttention layers use
rope_theta, partial_rotary_factor, and optional yarn. The optional
attention_heads_per_layer array overrides the base Q-head count for each
layer. Attention projection gating is tensor-presence driven:
self_attn.g_proj.weight [num_heads, hidden] means per-head
softplus(g_proj·x) gating immediately before o_proj (a
[num_heads·head_dim, hidden] projection means per-channel gating).
These fields and tensor semantics cover Laguna without introducing a
model-name-specific execution operator.
2.1 MTP — multi-token prediction (optional)
If the model carries an MTP head (DeepSeek/Qwen style), arch declares:
"mtp": { "num_layers": 1, "share_lm_head": true, "share_embed": true }
MTP tensors are ordinary directory entries under canonical names
(model.mtp.*): enorm.weight, hnorm.weight,
eh_proj.weight [hidden, 2·hidden], layers.{i}.* (a standard
transformer block), norm.weight.
Semantics: x = eh_proj·[enorm(embed(t_{p+1})); hnorm(h_p)] — embedding
FIRST (oracle-verified: the reverse order yields exactly 0% acceptance)
→ block → shared lm_head → draft of token t_{p+2}. A reader is not
required to execute MTP (metadata + ordinary tensors, additive
evolution, no feature bit); the CMF runtime uses the head for
speculative decode with a strict guarantee: output is exactly equal to
plain greedy — a rejected draft is rolled back from KV.
2.2 MoE — mixture-of-experts FFN (optional)
If the model carries MoE layers (Qwen2-MoE / Qwen3-MoE / Qwen3.5-MoE), arch declares:
"moe": {
"num_experts": 256, "top_k": 8, "moe_intermediate_size": 512,
"norm_topk_prob": true, // Qwen2-MoE: false
"shared_expert_intermediate_size": 512, // absent if no shared expert
"router_sigmoid": true, // optional; default = softmax
"routed_scaling_factor": 2.5 // optional; default = 1
}
Tensors are ordinary directory entries under HF names:
model.layers.{i}.mlp.gate.weight [num_experts, hidden] router
model.layers.{i}.mlp.experts.{e}.{gate,up,down}_proj.weight
model.layers.{i}.mlp.expert_bias [num_experts] selection only
model.layers.{i}.mlp.shared_expert.{gate,up,down}_proj.weight
model.layers.{i}.mlp.shared_expert_gate.weight [1, hidden] optional
Which layers are MoE is decided by the PRESENCE of the router in the
directory (per-layer, not per-model): Qwen2-MoE's
mlp_only_layers/decoder_sparse_step produce mixed models, and dense
layers keep ordinary mlp.*_proj.
Execution semantics (HF parity, gated by tests/moe_parity.sh across
multiple families): by default, softmax over ALL router logits; when
router_sigmoid, score each expert independently with sigmoid. An optional
expert_bias affects top-k selection only, not the gathered weights. Select
top-k (ties: lower index), optionally renormalize the selected weights, then
apply routed_scaling_factor and compute Σwₑ·FFNₑ(x). The shared expert is
always added: with weight sigmoid(shared_expert_gate·x) when that tensor is
present, otherwise with weight 1 (Laguna).
Experts stay quantized in mmap; per token only the pages of the selected
k are touched — the same residency story as skills. Writers SHOULD lay
a layer's expert tensors out role-contiguously (all gate_proj of
experts 0…N−1 back to back, then all up_proj, then all down_proj)
— GPU backends can then treat a layer's expert bank as one region
instead of gathering hundreds of slices; the native importer and
moe-defrag both emit this order. Each expert is a
separate directory entry with ITS OWN dtype: that is the carrier of
per-expert bit allocation (P15 claim 12) — implemented, gated by
tests/moe_vbit.sh; the B-field (router selection frequencies via
--route-stats) was measured end-to-end on a 35B model.
3. Tensor directory
Byte-for-byte the .vmfc v2 layout (single canon, shared reference
parser):
[0 : 8 ] count : u64
[8 : 16] pool_off : u64 (name-pool offset from section start)
[16 : 16 + count·56] 56-byte records:
name_off : u32 (relative to pool_off)
name_len : u16
dtype : u8 (§3.1)
ndim : u8 (≤ 6)
shape : u32 × 6 (zero-padded tail)
off : u64 (RELATIVE to data_off; multiple of 64)
nbytes : u64
hash : u64 (hash64 of the tensor bytes, §8)
[pool_off : …] UTF-8 name pool
Tensor names are 1:1 with the source model
(model.layers.{i}.mlp.gate_proj.weight, model.embed_tokens.weight,
lm_head.weight, …). The format does not prescribe a tensor set: the
directory is the single source of truth for what the blob contains.
There is no "computable layout".
3.1 dtype
Numbering shared with .vmfc (ids are never reused):
| id | name | status in CMF v2 |
|---|---|---|
| 0 | f32 |
✅ read/write |
| 1 | f16 |
✅ read/write (norms and 1-D are always f16) |
| 2 | bf16 |
✅ read/write |
| 3 | q8_row |
✅ read/write |
| 4 | q4_block |
✅ read/write |
| 5 | mix8_4 |
reserved |
| 6 | u8 |
reserved |
| 7 | q4_col |
reserved |
| 8 | vbit |
✅ read/write (QUANT_2F bit), variable 3–8 bit |
| 9 | q8_2f |
✅ read/write (QUANT_2F bit), 𝒲×θ |
| 10 | vbit_ro |
✅ read/write — vbit + in-file row-offset table (O(1) row access); converter default for --quant vbit |
| 11 | q4_tiled |
✅ read/write — q4 in interleaved [f16 scale][16B nibbles] tiles (--quant q4t) |
| 12 | q1 |
✅ read/write — 1-bit binary, for 1-bit-TRAINED models only (--quant q1) |
| 13 | q1s |
✅ read/write — q1 base + sparse high-precision outlier overlay (1-bit PTQ of normal checkpoints) |
| 14 | q1t |
✅ read/write — ternary {−s, 0, +s} base-3 tiles + per-row outlier overlay (~2.25 bpw + overlay) |
| 15 | q4tp |
✅ read/write — q4_tiled nibbles with the per-tile scale as a 5-bit rung on a per-row ladder (--quant q4tp, or requant in place) |
3.2 Quant layouts (canon = .vmfc: "quants first, then scales")
q8_row(2-D[out, in]only):[int8 : out·in][f16 : out]— one scale per row,w = q[o,i]·scale[o],scale[o] = absmax(row_o)/127.q4_block: groups of 32 over the flattened tensor, zero-padded;[u8 : ceil(n/32)·16][f16 : ceil(n/32)]. Nibbles: element2klow,2k+1high;w = (q − 8)·scale,scale = absmax(group)/7.- 1-D tensors and tensors < 32 elements are always
f16(normalization precision at maximal matrix compression). q8_2f:[int8][f16 row-scale][f16 col-field],w = q·scale[o]·col[i]— the two-field Madelung split 𝒲×θ, validated in vmfcore (+37% at equal size; recovers ~75% of the q8→f16 gap on outlier input channels).vbit(2-D only,in % 32 == 0; P13 FIG.3):[u8 bits: rows][f16 scales: rows·in/32][bit-packed rows, MSB-first, each row padded to a byte];w = (u − L)·scale[r,g],L = 2^{b−1}−1, levels b ∈ {3,4,5,6,8}, floor 3 (claim 13). Allocation b_r: water-filling over the log2 row amplitude toward the tensor's mean budget; for MoE experts the budget is SHARED across the family (layer × projection): the shiftā_expert − ā_familyis equivalent to joint water-filling over all experts' rows — a loud expert gets more bits, a quiet one is pinned to the floor (P15 claim 12; gatetests/moe_vbit.sh). Optionally the allocation takes the product with a B-field — router selection frequencies collected at calibration (b ∝ log2(A·B), truncated Fisher).vbit_ro(2-D only,in % 32 == 0): the same bits/scales/packed encoding asvbit, plusu32 row_offsets[rows+1](relative to the packed area) between the scales and the packed rows —[u8 bits: rows][f16 scales: rows·in/32][u32 offsets: rows+1][packed]. Readers get O(1) row access without a prefix scan over bit widths. The byte semantics ofvbit = 8are untouched; new id on purpose.q4_tiled(2-D only,in % 32 == 0):repeat per 32-group { [f16 scale][16B nibbles] }— 18-byte tiles, one sequential memory stream instead of two distant ones. Values and nibble order are identical toq4_block; only the placement of the scale differs (kernel-measured ×1.66 ARM / ×1.13 AVX2 over split).q4tp(2-D only,in % 32 == 0):[nibbles: rows·gpr·16][row params: rows × (f16 lo, f16 step)] [codes: rows × ceil(gpr·5/8), 5-bit LSB-first, row-aligned],gpr = in/32. A tile's scale is2^(lo[r] + code·step[r]), so a reader expands one row's 32-rung ladder once and then reads scales by table lookup. Nibble values and order are identical toq4_tiled; only the scale's representation differs. 4.17 bits/weight against 4.50 — the f16 scale was 11% of a q4t file.lo/stepcome from the row's exact min/max log-scale, so no code is ever out of range and the format needs no escape hatch. Encoders MUST roundlo/stepto f16 before choosing codes, and MUST quantize the nibbles against the reconstructed scale — otherwise writer and reader disagree, the same trap that makes a q4 encoder round its scale first.q1(2-D only,in % 32 == 0):repeat per 32-group { [f16 scale][4B sign bits] }— 6-byte tiles, 1.5 bits/weight. Bit k of byte j (LSB-first) is weight j·8+k of the group;w = scale·(2·bit − 1) ∈ {−s, +s},scale = mean|group|(the L2-optimal binary level). Intended for 1-bit-TRAINED models (Bonsai / BitNet class), where per-group weights already sit on two levels and the encoding is lossless up to f16; as post-training quantization of a normal checkpoint it destroys quality, so converters expose it only as an explicit opt-in.q1s(2-D only,in % 32 == 0): aq1base (identical 6-byte tiles; outliers are EXCLUDED from the group scale) followed by a sparse high-precision overlay:[u32 count]thencount × { [u32 flat-index][f16 value] }— the salient weights kept at full precision (holographic transfer / SpQR-style) and restored verbatim at dequant. Variable length:expected_nbytesis undefined, the reader trusts the directory's stored span. Lets a NORMAL checkpoint survive 1-bit where plainq1cannot.q1t(2-D only,in % 32 == 0,inmust fitu16): ternary BitNet-b1.58-style{−s, 0, +s}. Base:repeat per 32-group { [f16 scale][7B base-3 codes] }— 9-byte tiles, 5 ternary values per byte (3⁵ = 243 ≤ 256; code 0 → 0, 1 → +s, 2 → −s), ~2.25 bits/weight. Then a per-row outlier overlay:[u32 row_ptr[rows+1]]followed by{ [u16 col][f16 value] }entries grouped by row (rowr's outliers are[row_ptr[r], row_ptr[r+1]);colis a within-row index) — 4 bytes per outlier, no binary search. Capturing the many near-zero weights exactly is the decisive PTQ win over binary. Variable length, same span rule asq1s.
4. Weight blob
data_off is a multiple of 4096 (page-aligned mmap); every tensor
inside starts on a 64-byte boundary (SIMD loads, cache lines). Zero
padding between tensors. A reader interprets the blob only through the
directory.
5. Masks section
A task mask = bit fields of "what is active" over shared weights (weights do not change — the VMF principle: a skill selects a subset of the condensate).
[0 : 4] n_masks : u32
[4 : 8] meta_len : u32
[8 : 8 + meta_len] JSON meta (§5.1)
[…] mask blobs, each aligned to 8 from the section start
One mask blob (sizes derived from arch, no internal headers):
[n_layers × ffn_bytes] FFN bitfields ffn_bytes = ceil(intermediate_size / 8)
[n_layers × head_bytes] head bitfields head_bytes = ceil(num_attention_heads / 8)
[gates_bytes] layer_gates gates_bytes = ceil(num_layers / 8)
[n_layers × expert_bytes] expert bitfields OPTIONAL — only when the mask's meta
sets "has_expert_fields": true;
expert_bytes = ceil(moe.num_experts / 8)
Bit order is LSB-first: neuron i = bit i % 8 of byte i / 8; bit
set → active. Tail bits beyond the dimension MUST be zero (or
popcount sees phantom neurons/heads).
The optional expert area (additive: old readers never look past the
gates, and each mask's blob_len is explicit) makes a task mask narrow
MoE ROUTING: bit e of layer l's row set → expert e is routable
for this task; selection then happens over the routable set only, the
router softmax renormalizing over it. This is the runtime-switchable
twin of §11.1's physical expert defrag — one file with the full expert
set serves many specialists (cortiq moe-mask writes such masks,
run --task <name> activates one; verified token-identical to the
equivalent runtime restriction). A layer whose row is all-ones is
unrestricted; a mask without the area restricts nothing.
5.1 Mask JSON meta
{
"default_task": "general",
"masks": [{
"task_id": 0, "name": "general", "description": null,
"sparsity": 0.62,
"quality": { // null = NOT MEASURED (declaring 1.0 is forbidden)
"metric": "heldout_ppl_ratio", "value": 0.97,
"baseline_dense": 6.10, "n_samples": 512, "dataset_sha256": "…"
},
"parent": null, "priority": "Fallback", "has_hot_pack": false,
"blob_off": 4096, "blob_len": 139328 // relative to section start
}]
}
quality is a held-out contract, not a declaration: a converter
without a measured metric writes null; the runtime logs a warning when
switching to an unmeasured mask.
6. Tokenizer section
The bytes of HuggingFace tokenizer.json, verbatim. The model is
self-contained: one file = one unit of distribution. A sidecar file
remains a debugging fallback.
6.1 Chat bundle (header.tokenizer_config)
The file — not the runtime binary — defines chat behavior. The header carries an optional block (additive evolution, no feature bit):
"tokenizer_config": {
"chat_template": "<Jinja template from chat_template.jinja or tokenizer_config.json>",
"eos_token_ids": [248044, 248045],
"bos_token_id": null,
"pad_token_id": 248055
}
The runtime renders the template with HF semantics (trim_blocks,
lstrip_blocks, loop controls, Python string methods) and stops
generation on any id in eos_token_ids. Gate:
tests/chat_template_parity.sh — the runtime render equals reference
jinja2 byte-for-byte. Files without the block get a ChatML fallback.
7. Sparse index
A precomputed bridge "mask → computation skip": active FFN quant groups (32 neurons each) and heads, per (task, layer) pair.
Honest status: the engine takes active indices directly from the mask bitfields; the index is read and displayed by the CLI but has never been used in execution. Deprecation-pending: writers SHOULD stop emitting it (readers keep parsing existing files); it is revived only if the "masks × quantized mmap" path materializes with a measured win.
[0 : 4] n_entries : u32
[4 : 8] reserved : u32 (0)
entry (4-aligned):
task_id : u32
layer_idx : u32
n_groups : u32
n_heads : u32
[u16 × n_groups] active FFN-group indices (sorted)
[u8 × n_heads] active head indices (sorted)
zero padding to a multiple of 4
A group is active if it contains at least one active mask bit.
8. hash64
A non-cryptographic 64-bit hash of tensor bytes: murmur3 fmix64 over
64-bit LE words with positional salt i·0x9E3779B97F4A7C15, XOR fold,
xor len, final fmix64. Bit-for-bit compatible with
vmfcore.hash64 (Python) and vmfcore::hash64 (Rust) — hashes of
shared tensors match between .cmf and .vmfc (backbone dedup across
skill files is free).
Uses: cortiq verify (corruption detection), dedup, cache keys.
8.1 Section hashes
Metadata integrity (not just tensors):
- Envelope reserve
[0x70:0x78]= hash64(header JSON),[0x78:0x80]= hash64(directory). Zero = "absent" (older files pass). - The header JSON carries
section_hashes— hex hash64 of masks/vocab/index (u64 as a JSON number would lose precision past 2^53). The header hash in the envelope transitively covers them. - The envelope itself (first 0x70 bytes) is not hashed: a hash cannot protect itself; corrupted offsets are caught by bounds/hashes further down the chain.
cortiq verifychecks the whole chain; a single flipped header byte is an error.
8.2 Detached signature (authenticity, opt-in)
The hash chain proves integrity, not authorship. cortiq sign writes a
detached <model>.sig — JSON {alg: "ed25519-sha256", pubkey, sha256, sig}, Ed25519 over the file's SHA-256 — so the container itself is
never rewritten and old tooling is untouched. cortiq verify checks
the signature automatically when the .sig sits next to the model;
absence is not an error. Key = a 32-byte hex seed file the signer
keeps private.
Anti-features — what the format deliberately does NOT have
- A computable weight layout — bug class #1 of v1 (writer and reader "computed" the layout independently and diverged).
- Silent fallbacks — v1 would interpret any garbage file as "a 27B model"; v2 must fail.
- JSON for bit data — v1 masks in JSON bloated 3–4×.
- Declaration fields —
quality_score: 1.0by default, area-law "capacities", Born multipliers in dynamics: a metaphor does not become a format field until it is measured.
9. Skills — a swarm in one file (Patent 15, claims 2/12/15)
One shared backbone + K per-skill records; no record stores a full model. Storage scales as |backbone| + Σ|deltas|.
Replacement tensors are ordinary directory entries named
skill.{skill_id}.{name_of_replaced_tensor}, e.g.
skill.sql.model.layers.3.mlp.gate_proj.weight. The full logical shape
of the replaced tensor (full-shape — NOT low-rank, NOT a diff list, NOT
a mask), in any encoding of §3. The per-skill delta index (claim 2) is
materialized by the directory: a prefix filter yields skill →
byte-offsets; lazy paging = mmap access to exactly those offsets
(claim 12).
Registry — header JSON, additive:
"skills": [{
"id": "sql",
"name": "SQL assistant",
"layers": [3, 4, 5],
"selection": {"metric": "mse", "phi_layer": 20,
"mean": "<f16 base64>", "basis": "<f16 base64>"},
"input_mask_task": null,
"quality": {"metric": "ppl", "backbone": 21.4, "overlaid": 17.9,
"dataset_sha256": "…"}
}]
selection holds the affine-subspace parameters for recon-argmin
routing (E = ‖r − BBᵀr‖²/‖φ‖², choose the skill with minimal E); the
file is self-sufficient for selection. quality is the honest claim-16
contract (overlaid vs backbone on held-out data).
Execution semantics (claims 1/3/18): tensor-source indirection — for
every tensor the runtime reads EITHER the backbone entry OR
skill.{active}.{name} if present; replacement instead of addition, a
full per-skill model is never assembled (all tensors are pointers into
one mmap). Soft superposition (claim 14): blended working tensors
Σwᵢ·Tᵢ, wᵢ = softmax(−E/T).
Append-only growth (claim 11): adding a skill = appending new
tensors at the file tail + re-emitting directory/header/index at the
tail + updating envelope offsets in place (offset 0 is fixed). Bytes and
offsets of previously written tensors never change; old dir/header bytes
become dead section tails (compatible: readers navigate only through the
envelope). Compaction (converter/cmf_compact.py) = a plain rewrite.
9.1 Standalone skill files (SKILL_FILE, bit 6)
A skill can also travel WITHOUT its backbone: a .cmf whose tensor set
is only what a bake changed (plus the mask catalog), bound to the base
it was cut against by identity keys in the registry record:
"skills": [{
"id": "gfx-html",
"layers": [0, 1, "...", 21],
"base_dir_hash": "9f22593eb458bc6f",
"base_arch": "nanbeige",
"task": "specialist",
"provenance": {"corpus": "…", "tensors": 30}
}]
base_dir_hash— hexhash64of the BASE file's tensor-directory bytes (the same value the envelope carries at[0x78]). A skill is a delta against exact bytes, not against an architecture:applyMUST refuse a base whose directory hash differs (an explicit--forcemay override; the result is out of spec).base_arch,task,provenance— informative keys: the human check, the mask-catalog task the skill activates, and where it came from.
Any record with base_dir_hash present raises feature bit 6, so a
pre-bit reader refuses the file loudly and a runtime that knows the bit
refuses to RUN it (a partial tensor set is not a model) and points to
cortiq skill apply <base> <skill> -o out.cmf, which verifies the key,
overlays tensors and masks over the base, and writes a complete file —
byte-equivalent to the specialist the skill was cut from.
Lifecycle: skill bake (specialist) → skill export --base (delta +
keys) → publish the small file → skill apply on any copy of the base.
Status: fully implemented and gated (container + indirection, production recipes, recon-argmin routing, append-only + compaction, soft-blend); claim 16 met by measurement (−24.9% task-PPL in the runtime).
10. Sharding — a model in N files
Naming: {base}-{no:05}-of-{count:05}.cmf (spiritually compatible with
safetensors). The user opens ANY name; the runtime normalizes to shard 1
and picks up siblings by pattern.
Every shard is a standalone valid .cmf: full envelope, header JSON,
a directory of ITS OWN tensors, its own data blob, its own hashes
(section_hashes + per-tensor). cortiq verify works on any single
shard without its siblings.
Each shard's header carries:
"shard": { "no": 1, "count": 5 }
No block = an ordinary single file (backward compatible: old readers see shard 1 as a valid but incomplete model and fail honestly on the missing tensor).
Content distribution: tensors are split greedily in canonical order
(--shard-max-gb threshold, rough f32 size); the masks/vocab/sparse
index sections, tokenizer_config (chat bundle) and the skills
registry live ONLY in shard 1 — the rest have empty sections and
tokenizer_config: null. Skill tensors (skill.{id}.*) are distributed
as ordinary directory entries — the shard-1 registry references them by
name through the merged directory.
Loading (CmfModel::open_sharded): open shard 1 → mmap all siblings
→ merge directories (each entry remembers its shard index — a runtime
field, never written to disk) → the runtime then works as with a single
file. Errors: opening a non-first shard directly, a missing sibling, a
count mismatch.
Gate (Qwen3.5-0.8B q8_2f, 5 shards ≤ 0.6 GB): sharded PPL == unsharded
byte-exactly on the same binary; verify green on every shard alone.
11. Defragmentation — physical pruning (USPTO App. 19/452,464, claims 9/10 — PATENTS.md)
A mask (§5) is virtual sparsity: pruned neurons are flagged but still
stored in full (all tasks share one backbone — you cannot physically cut
it until you commit to ONE task). Defragmentation turns virtual sparsity
into physical compression: pruned FFN neurons are dropped from the
file — they are neither stored nor computed. This is Factory-Hard →
defrag from the DTG-MA application (19/452,464): "bake one mask into the weights" and emit a
standalone compact .cmf.
Representation — no new feature bit, backward compatible. Physical
pruning is expressed ONLY by smaller tensor shapes in the directory (§3
"no computable layout"; the directory is the sole shape authority). The
runtime derives the FFN size from the tensor shape (gate_proj.rows()),
not from arch.intermediate_size, so a defragged file is an ordinary
smaller dense model that existing readers load unchanged. The masks
section (§5) is absent in a defragged file (the mask is the identity
after pruning). arch.intermediate_size becomes nominal (= the per-layer
max); the true size lives in each tensor.
Per-layer variance — better than the patent. Because the directory
carries an arbitrary shape per tensor, each layer shrinks to its OWN
live-neuron count. The patent must truncate every layer to max(active)
(one bottleneck layer caps the ratio at 80.2% vs. 94% achievable) — CMF
has no such limit.
Invariants (mandatory):
- per-layer triple:
gate_proj.rows() == up_proj.rows() == down_proj.cols() == inter'ₗ, anddown_proj.rows() == hidden_size. One keep-set indexes all three (gate row i, up row i, down col i are the same neuron); - neuron axis: rows (axis 0) for
gate_proj/up_proj, columns (axis 1) fordown_proj; - quant group of 32: the
down_projneuron axis is its COLUMNS, andvbit/q4_blockrequirein % 32 == 0. Adown_projwhoseinter'is not a multiple of 32 is written asq8_2f(per-row scale — no column constraint; the converter downgrades automatically).gate/updrop rows, so their columns (= hidden) are unaffected; - NOT a byte truncation: quant scales are per-group/per-row, so pruning
is dequant → gather live neurons → requant at the smaller shape
(the
q8_2fcol-field /vbitscales ofdown_projregenerate for the shrunk column set); tensor hashes are recomputed; hidden_size,embed_tokens,lm_head, and norms are untouched (skill-selection subspaces depend on hidden).
One task, standalone file. Defrag is destructive: one .cmf bakes
exactly one task. Multi-task serving stays on masks (§5) or per-skill
replacement tensors (§9).
Provenance (honest contract). Header provenance.defrag:
"defrag": {
"source_skill": "…/skill_ru",
"pre_intermediate": 3072,
"post_intermediate_max": 640,
"kept_per_layer": [608, 640, 512, ...],
"pruned_ratio": 0.803
}
Numerically the dense output of a defragged model is IDENTICAL to the
masked output before quantization (a dead neuron contributes act·0
under a mask and is simply absent after defrag); after quantization the
only difference comes from quantizing the smaller matrices.
Scope: dense FFN neurons here; MoE experts in §11.1. Attention-head pruning (the head count is a global runtime scalar) is out of scope.
11.1 MoE expert defrag (cortiq moe-defrag)
The MoE twin of §11, driven by the routing B-field instead of a neuron
mask: expert usage is strongly task-conditional (measured on a 34.7B
coder: the top-64 expert sets for code vs prose overlap with Jaccard
0.25), so a one-task file can drop the experts that task never routes
to. From a CMF_MOE_STATS dump (per-layer expert-selection counts over
a task-representative run), keep per layer the smallest top expert set
reaching --cover of the recorded routing mass; drop the rest.
Representation — same philosophy as §11, no feature bit.
- Kept experts are renumbered into a CONTIGUOUS per-layer prefix
mlp.experts.0 … mlp.experts.{k−1}preserving relative order; a reader enumerates a layer's experts by tensor PRESENCE up toarch.moe.num_experts, which becomes nominal (= the original count) — mirroring §11's rule forintermediate_size. - The router tensor's rows are gathered to match, in the same order:
mlp.gate.weightbecomes[kept_l, hidden], androuter.rows() == (number of expert entries present)is a load-time invariant.top_kclamps to the per-layer expert count. - Selection semantics are unchanged (§2.2): the softmax simply
renormalizes over the kept set. The identical restriction can be
applied at RUNTIME without rewriting the file
(
CMF_MOE_MASK=<stats.json>+CMF_MOE_MASK_COVER) — the two are mathematically equal, which is how a cover level is perplexity-gated before committing to the cut. - Expert payloads are copied verbatim (no requant — the expert axis is whole tensors, not quant groups), so the surviving weights are byte-identical to the source and the rewrite streams from the source mmap.
Measured reference (KAT-Coder 34.7B-A3B, code-calibrated, cover 0.95): 19.6 → 12.7 GB (−35%), held-out code perplexity +2.8%, and on a 24 GB machine — where the full model paged — decode ×1.8, prefill ×3.3.
Off-task quality degrades by design; like §11, one defragged file bakes one task. Multi-task serving stays on the full expert set.
Producing it (native Rust):
cortiq convert --model <hf_dir_or_repo> --defrag <skill_dir> \
--quant q8_2f --output model.cmf
<skill_dir> carries baked FFN overlays (tensors/*.npy) and, if
available, a keep-set ffn_keep.npy (bool [n_layers, intermediate],
True = live) from the pruning pipeline. Without ffn_keep.npy the
keep-set is autodetected from all-zero down_proj columns (the
Factory-Hard bake). The mask-training / bake step lives in the private
research pipeline; the public tool only consumes its artifacts.
12. Pipeline containers — text-to-image in one file
The same envelope/directory/blob machinery carries non-LLM pipelines.
The only differences are the arch_name tag and namespaced tensor
names; no new sections, no feature bit (a reader that does not execute
the pipeline still validates and inspects the file).
Current instance — arch_name: "lumina2-image" (Lumina-Image 2.0,
cortiq imagine-pack / cortiq imagine): one file packs the whole
text-to-image stack.
- Namespaces:
te.*— the text-encoder transformer (a Gemma-2 class LLM; the header'sarchblock describes THIS component, so generic tooling reads meaningful dimensions),dit.*— the Next-DiT denoiser,vae.*— the VAE decoder. Component config JSONs ride as{prefix}.config_jsonu8 tensors — the file is self-sufficient. - Quantization: per-tensor as always (§3 directory is the truth) — typically q4t/q8 matrices for te/dit, f16 for VAE convolutions and norms.
- Tokenizer section (§6) carries the text encoder's tokenizer;
provenance.pipeline+provenance.componentsname the recipe.
The measured reference lives in the README (512 px on CPU in minutes, Metal whole-DiT-block graph on Apple silicon; the wgpu path serves discrete cards and phones).
Related: COMPARISON.md (CMF vs. other model formats),
project README (overview and quick start),
python/cmf_reader.py (standalone reader: stdlib + numpy, reads every
dtype, shards, skills, verify).