# CPU and runtime status ## What works now `checkpoint/loader.py` is self-contained and supports a PyTorch device argument, including `cpu`. It verifies the checkpoint, creates the Qwen3 architecture, decodes the T3+sparse base, applies WALB2 overlays, restores initialized RoPE buffers exactly, and returns an evaluation model. No dense Qwen or Neutrino tensor file is opened during this process. ## Verified release CPU smoke test The exact v0.1 package was tested offline on the project's many-core server: - command: `python checkpoint/loader.py smoke checkpoint --device cpu`; - checkpoint hashes verified before loading; - total load time: 348.36 seconds; - one five-token forward pass: success; - logits: shape `[1, 151936]`, all finite; - peak resident set: 16.08 GiB; - dense Qwen tensor files opened: no; - donor or Neutrino tensors opened: no. This proves functional CPU compatibility for the reference loader. It does not establish production tokens/s or portability of that timing to a different CPU. The machine-readable record is `evidence/cpu_smoke_v0_1.json`. ## What is not implemented yet The loader materializes every parameter as BF16. It does not execute matmuls directly over packed ternary/sparse/binary structures. Consequently: - serialized disk size is about 2.95 GB; - resident parameters are approximately 16.4 GB in BF16; - additional working memory is required while decoding and generating; - CPU inference will use ordinary PyTorch BF16/FP32 kernels and may be slow; - this is not equivalent to a llama.cpp/BitNet.cpp deployment. At least 32 GB system RAM is recommended for CPU experiments. More is safer for long contexts. CPU throughput is not yet a release claim. ## Reference GPU measurements On the project's physical GPU2: - verified V73 load: 374–377 seconds; - C4 evaluation: about 21,772 scored tokens/s at batch 2×2048; - MMLU evaluation after load: 87.68 seconds; - GSM8K first-300 generation evaluation: 436.53 seconds. These are protocol-specific research timings, not production serving numbers. ## Next runtime milestone Implement direct packed kernels or export into a runtime that can express: 1. radix-3 ternary plane; 2. exactly-k8 signed sparse lane per group128; 3. binary low-rank correction paths; 4. INT3 embedding and INT4 head. Only after that should CPU resident memory and tokens/s be advertised as a low-bit runtime result.