| # CPU and runtime status |
|
|
| ## What works now |
|
|
| `checkpoint/loader.py` is self-contained and supports a PyTorch device argument, |
| including `cpu`. It verifies the checkpoint, creates the Qwen3 architecture, |
| decodes the T3+sparse base, applies WALB2 overlays, restores initialized RoPE |
| buffers exactly, and returns an evaluation model. |
|
|
| No dense Qwen or Neutrino tensor file is opened during this process. |
|
|
| ## Verified release CPU smoke test |
|
|
| The exact v0.1 package was tested offline on the project's many-core server: |
|
|
| - command: `python checkpoint/loader.py smoke checkpoint --device cpu`; |
| - checkpoint hashes verified before loading; |
| - total load time: 348.36 seconds; |
| - one five-token forward pass: success; |
| - logits: shape `[1, 151936]`, all finite; |
| - peak resident set: 16.08 GiB; |
| - dense Qwen tensor files opened: no; |
| - donor or Neutrino tensors opened: no. |
|
|
| This proves functional CPU compatibility for the reference loader. It does |
| not establish production tokens/s or portability of that timing to a |
| different CPU. The machine-readable record is |
| `evidence/cpu_smoke_v0_1.json`. |
|
|
| ## What is not implemented yet |
|
|
| The loader materializes every parameter as BF16. It does not execute matmuls |
| directly over packed ternary/sparse/binary structures. Consequently: |
|
|
| - serialized disk size is about 2.95 GB; |
| - resident parameters are approximately 16.4 GB in BF16; |
| - additional working memory is required while decoding and generating; |
| - CPU inference will use ordinary PyTorch BF16/FP32 kernels and may be slow; |
| - this is not equivalent to a llama.cpp/BitNet.cpp deployment. |
|
|
| At least 32 GB system RAM is recommended for CPU experiments. More is safer for |
| long contexts. CPU throughput is not yet a release claim. |
|
|
| ## Reference GPU measurements |
|
|
| On the project's physical GPU2: |
|
|
| - verified V73 load: 374–377 seconds; |
| - C4 evaluation: about 21,772 scored tokens/s at batch 2×2048; |
| - MMLU evaluation after load: 87.68 seconds; |
| - GSM8K first-300 generation evaluation: 436.53 seconds. |
|
|
| These are protocol-specific research timings, not production serving numbers. |
|
|
| ## Next runtime milestone |
|
|
| Implement direct packed kernels or export into a runtime that can express: |
|
|
| 1. radix-3 ternary plane; |
| 2. exactly-k8 signed sparse lane per group128; |
| 3. binary low-rank correction paths; |
| 4. INT3 embedding and INT4 head. |
|
|
| Only after that should CPU resident memory and tokens/s be advertised as a |
| low-bit runtime result. |
|
|