File size: 21,355 Bytes
9e551a1 be4e085 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 | ---
license: agpl-3.0
tags:
- snapkitty
- october-2026-drop
- python
- cuda-q
- qsharp
- quantization
- distillation
- model-compression
- machine-learning
---
> Mirrored from [https://github.com/SNAPKITTYWEST/tensor-roll](https://github.com/SNAPKITTYWEST/tensor-roll) at commit `f8cb763`. Part of the **SnapKitty October 2026 main drop**.
# TENSOR ROLL v1.0
**Recursive CUDA-Q Model Quantizer β compress a 30B-class teacher into an ~8B student under a hard parameter budget, from the terminal.**
```
GLIMMER 30B β TENSOR ROLL β CUDA β CUDA-Q β Q# β 8B NANO MODEL
```

Tensor Roll treats model compression as a searchable computational system. A recursive
`TensorRoll(T, axis, depth, rank, quantum_policy)` operator partitions every tensor,
measures each partition (norm, rank, entropy, spectral contribution, reconstruction error),
and searches a global representation plan under a hard budget of **P β€ 8.5B parameters**
(target ~8.0B). The resulting student is distilled with a 7-term loss, fine-tuned, exported
to FP16 / BF16 / INT8 / INT4 and a checksummed `.trq` container β then served from a
sandboxed terminal runtime.
Everything is written from scratch on NumPy: the tensor core, the transformer,
distillation, the quantizers, the `.trq` codec, the OS sandbox, and the CLI.
Every execution report carries an honest backend label β `cpu-numpy`, `qsim-classical`,
or `unavailable` β so a backend is never claimed that wasn't actually executed.
## Features
- **Recursive TensorRoll operator** β `TensorRoll(T, axis, depth, rank, quantum_policy)`
recursively partitions tensors and evaluates eight decision classes per region
(`PRESERVE`, `QUANTIZE`, `FACTORIZE`, `MERGED`, `ROUTED`, `RECONSTRUCTED`,
`PRUNED`, `QUANTUM-ENCODED`) from measured (cost, retained-energy, error) triples.
- **Out-of-core streaming (Rule Zero)** β the teacher is never required in RAM.
Sharded, memory-mapped ingestion with a configurable working-set ceiling
(`--memory-budget`), recursive tensor windowing, explicit eviction, and
checkpoint/resume.
- **SCAN β PLAN β EXECUTE global budget planner** β lightweight scan, global
allocation solved under Ξ£ target β€ 8.5B, then streaming execution.
Sensitivity-aware: budget flows non-uniformly to the tensors that earn it.
- **From-scratch NumPy core** β no PyTorch, Transformers, llama.cpp, or ONNX
wrappers. The only runtime dependency is NumPy.
- **7-term recursive distillation** β task, logits, hidden-state, attention,
embedding, reconstruction, and roll losses, with resumable stages 0β7.
- **FP16 / BF16 / INT8 / INT4 + `.trq`** β multiple physical exports plus the
Tensor Roll container format: magic header, tensor index, quantization
metadata, codebooks, per-tensor and file-level SHA-256 checksums, provenance.
- **OS-level sandbox** β process-group isolation, resource limits, environment
filtering, jail path policy, allow/deny binary policy, network namespace
isolation, and a JSONL audit log. Model tool requests route through the
sandbox supervisor.
- **14-command terminal CLI** β one coherent surface (`tensor-roll`) from
ingestion to chat, plus `tensor-roll doctor` for the backend capability report.
- **Honest backend dispatch** β classical GEMM, CUDA-Q kernels, and Q# operations
share one dispatch boundary with explicit per-invocation reporting
(backend, device, dims, precision, time, memory, depth, shots, reconstruction
error). `tensor-roll doctor` distinguishes *source present* from
*backend executed*.

## A novel method, executed
Tensor Roll is not a compression proposal β it is a compression method that
runs. What is novel is the composition, and every piece of it exists as code
in this repository:
**1. The recursive `TensorRoll` operator is the compression primitive.**
Most quantizers apply one fixed recipe (prune, quantize, hope).
`TensorRoll(T, axis, depth, rank, quantum_policy)` β implemented in
`tensor_roll/core.py` β recursively partitions every tensor, measures each
partition (norm, SVD rank, entropy, variance, spectral contribution,
reconstruction error), and *searches* a representation per region from eight
decision classes: `PRESERVE`, `QUANTIZE`, `FACTORIZE`, `MERGED`, `ROUTED`,
`RECONSTRUCTED`, `PRUNED`, `QUANTUM-ENCODED`. The decision is data-driven,
per region, per recursion level β executed by `tensor-roll roll`.

**2. Rule Zero: the model is never required in RAM.**
`tensor_roll/ooc.py` streams the teacher through sharded, memory-mapped
windows with a hard working-set ceiling (`--memory-budget`), explicit
eviction, and crash-safe checkpoint/resume. A 30.19B-parameter manifest is
planned to an 8.000B target in under 0.1 s at 35 MB peak RSS with zero
weights allocated β the planner reasons over metadata, not tensors.
**3. SCAN β PLAN β EXECUTE solves the budget globally.**
`tensor_roll/planner.py` scans cheaply, then allocates the β€ 8.5B budget
across all tensors with a sensitivity-aware greedy knapsack over measured
(cost, retained-energy) pairs β budget flows non-uniformly to the tensors
that earn it β before streaming execution begins. One global plan, not
greedy layer-by-layer heuristics.
**4. Heterogeneous dispatch with honest backend states.**
`tensor_roll/boundary.py` lowers one backend-neutral TensorRoll IR
(`LOAD Β· ROLL Β· GEMM Β· QUANTIZE Β· EMIT`, `tensor_roll/ir.py`) to CPU-NumPy,
CUDA, CUDA-Q, and Q# (`cuda/`, `cudaq/`, `qsharp/`). `tensor-roll doctor`
(`tensor_roll/doctor.py`) probes every backend at runtime and reports
`AVAILABLE` / `UNAVAILABLE` / `NOT EXECUTED` per invocation β source present
is never reported as backend executed.

**5. The `.trq` artifact is versioned, checksummed, and self-describing.**
`tensor_roll/trq.py` writes `TRQ1` containers: JSON header, per-tensor index
with dtype/quant codes, codebooks, per-tensor and file-level SHA-256, and
provenance records. Corrupt or tampered artifacts are rejected on load β
verified by round-trip tests across all five variants
(FP32 / FP16 / BF16 / INT8 / INT4).
Executed means: 70/70 tests green, 14/14 CLI commands live, and every number
in the [Measured benchmarks](#measured-benchmarks) gallery below was produced
by running this code on a real machine.
## Requirements
Python 3.10+ and NumPy. CPU + NumPy runs everywhere; CUDA, CUDA-Q, and Q#
backends activate automatically when the hardware/toolchain is present β
`tensor-roll doctor` reports exactly what is available on your machine.
## Installation
```bash
git clone https://github.com/AHMADALIPARR/tensor-roll.git
cd tensor-roll
pip install .
```
This installs the `tensor-roll` console script (`pyproject.toml`, setuptools).
For development, `pip install -e .` works the same way.
Verify:
```bash
$ tensor-roll --help
usage: tensor-roll [-h] [--workdir WORKDIR]
{inspect,ingest,map,roll,qkernel,compress,train,finetune,quantize,evaluate,benchmark,chat,sandbox,demo}
...
Tensor Roll β Recursive CUDA-Q Model Quantizer
```
## Quickstart

Check your backends, then run a tiny end-to-end pass:
```bash
$ tensor-roll doctor
Tensor Roll Backend Report
CPU AVAILABLE
NumPy AVAILABLE
CUDA UNAVAILABLE
CUDA-Q UNAVAILABLE
Q# UNAVAILABLE
GPU NONE
CUDA source VERIFIED
CUDA-Q source VERIFIED
Q# source VERIFIED
Execution:
CPU PASS
CUDA NOT EXECUTED
CUDA-Q NOT EXECUTED
Q# NOT EXECUTED
```
```bash
$ tensor-roll --workdir ./tr-quick ingest --scale tiny --steps 20
training tiny teacher Config(vocab=32, d=48, heads=3, layers=2, dff=96, seq=16) (41.1K params, seed=11)
teacher: 41.1K params, held-out loss=3.0724 acc=0.149, 3.9s on cpu-numpy
$ tensor-roll --workdir ./tr-quick map
tensor shape params rank entropy spectral
L0.W1 [48, 96] 4608 48 3.632 0.055
L0.W2 [96, 48] 4608 48 3.607 0.059
L0.Wk [48, 48] 2304 48 3.377 0.077
...
$ tensor-roll --workdir ./tr-quick roll --depth 2
[ROLL 0001] source: L0.W1/0/0 shape: [12, 96] depth: 2 params: 1152 rank: 12 entropy: 2.434 spectral: 0.125 cost: 288 retained-energy: 1.0000 recon-err: 4.60e-03 backend: cpu-numpy decision: QUANTIZE
[ROLL 0002] source: L0.W1/0/1 shape: [12, 96] depth: 2 params: 1152 rank: 12 entropy: 2.438 spectral: 0.123 cost: 288 retained-energy: 1.0000 recon-err: 4.61e-03 backend: cpu-numpy decision: QUANTIZE
...
```
The full demonstration runs the complete pipeline and drops into an inference REPL:
```bash
$ tensor-roll demo --fast
```

---
## User Guide
Every command accepts a global `--workdir DIR` (default: `./tensor-roll-work`).
All examples below use real flags from `tensor-roll <cmd> --help`.
### `inspect` β tensor inventory and metrics
Lists every tensor with shape, dtype, and Frobenius norm.
Flags: `--what {teacher,student,trq}`, `--trq TRQ`
```bash
$ tensor-roll --workdir ./tr-quick inspect --what teacher
teacher: Config(vocab=32, d=48, heads=3, layers=2, dff=96, seq=16), 41.1K params
L0.W1 [48, 96] float32 ||.||=8.1667
L0.W2 [96, 48] float32 ||.||=8.1073
L0.Wk [48, 48] float32 ||.||=6.9597
...
```
### `ingest` β build or load the teacher model
Trains a tiny from-scratch teacher (`--scale tiny`) or streams a large
teacher from sharded weights (`--scale 30b` with `--teacher` on `compress`).
Flags: `--scale {tiny,30b}`, `--steps`, `--batch`, `--lr`, `--seed`
```bash
$ tensor-roll --workdir ./tr-quick ingest --scale tiny --steps 20
training tiny teacher Config(vocab=32, d=48, heads=3, layers=2, dff=96, seq=16) (41.1K params, seed=11)
teacher: 41.1K params, held-out loss=3.0724 acc=0.149, 3.9s on cpu-numpy
```
### `map` β per-tensor metric map
One row per tensor: parameter count, numerical rank, entropy, and spectral
contribution β the raw material the planner allocates budget from.
```bash
$ tensor-roll --workdir ./tr-quick map
tensor shape params rank entropy spectral
L0.W1 [48, 96] 4608 48 3.632 0.055
L0.W2 [96, 48] 4608 48 3.607 0.059
...
```
### `roll` β run the TensorRoll operator
Recursively partitions each tensor and prints one `[ROLL n]` line per leaf:
source, shape, depth, measured metrics, backend, and the selected decision.
Flags: `--depth`, `--quantum-policy {off,explore,selective}`, `--budget-ratio`
```bash
$ tensor-roll --workdir ./tr-quick roll --depth 2 --quantum-policy explore
[ROLL 0001] source: L0.W1/0/0 shape: [12, 96] depth: 2 params: 1152 rank: 12 entropy: 2.434 spectral: 0.125 cost: 288 retained-energy: 1.0000 recon-err: 4.60e-03 backend: cpu-numpy decision: QUANTIZE
...
```
### `qkernel` β the tensor_roll kernel family
Runs `tensor_roll_encode / rotate / entangle / measure / reconstruct /
recursive` on a tensor block through the quantum path, with full
per-invocation reporting and a checksummed boundary trace.
Flags: `--tensor`, `--shots`, `--encoding {amplitude,angle}`, `--seed`
```bash
$ tensor-roll --workdir ./tr-quick qkernel --shots 512
tensor_roll kernel family on L0.W1(8, 8) (top-left 8x8 block, 64 elems -> 6 qubits, amplitude encoding)
tensor_roll_encode backend=qsim-classical device=cpu (statevector simulator) dims=[(8, 8), (64,)] prec=complex128 t=0.10ms mem=1.50KB depth=0 shots=β recon_err=β
tensor_roll_measure backend=qsim-classical device=cpu (statevector simulator) dims=[(64,)] prec=complex128 t=0.85ms mem=1.00KB depth=0 shots=512 outcomes=51 H=5.17b recon_err=β
tensor_roll_reconstruct backend=qsim-classical device=cpu (statevector simulator) dims=[(64,), (8, 8)] prec=float64 t=0.05ms mem=512.00B depth=0 shots=512 outcomes=51 H=5.17b recon_err=7.85e-01
...
block rel recon error: 0.7850 (recursive: 0.1916) in 0.00s β backend qsim-classical (statevector simulator, NOT a QPU)
boundary trace: /tmp/trqs/qkernel/boundary.json (28 gates, 54 outcomes, sha256 verified, recon[0]=0.1531)
```
### `compress` β search the 8B representation
The core command. Runs SCAN β PLAN β EXECUTE: scans tensor statistics,
solves the global budget (Ξ£ target β€ 8.5B, default target 8B), then streams
the transformation. `--out-of-core` enforces Rule Zero β the teacher is
memory-mapped and windowed, never fully loaded. `--resume` continues from
per-tensor checkpoints after an interruption.
Flags: `--budget-ratio`, `--seed`, `--teacher`, `--target-params`,
`--memory-budget`, `--out-of-core`, `--resume`
```bash
# In-RAM path (small teachers)
$ tensor-roll --workdir ./tr-quick compress
# Streaming path (large teachers) β bounded 5G working set, resumable
$ tensor-roll compress --teacher /models/glimmer-30b \
--target-params 8B --memory-budget 5G --out-of-core
$ tensor-roll compress --out-of-core --resume
```
### `train` β recursive distillation stages
Distills the student against the teacher with the 7-term loss
(task + logits + hidden + attention + embedding + reconstruction + roll).
`--stage all` runs stages 0β7; any single stage can be re-run. Checkpoints
make training resumable and deterministic.
Flags: `--stage` (`all` or 0β7), `--steps`, `--batch`, `--lr`, `--seed`
```bash
$ tensor-roll --workdir ./tr-quick train --stage all --steps 40
$ tensor-roll --workdir ./tr-quick train --stage 3 --steps 40 # re-run one stage
```
### `finetune` β task fine-tuning
Task fine-tuning pass on the distilled student.
Flags: `--steps`, `--batch`, `--seed`
```bash
$ tensor-roll --workdir ./tr-quick finetune --steps 40
```
### `quantize` β FP16 / BF16 / INT8 / INT4 + `.trq` export
Converts the trained student into every physical representation and writes
the `.trq` containers with quantization metadata, codebooks, scales,
checksums, and provenance.
```bash
$ tensor-roll --workdir ./tr-quick quantize
```
### `evaluate` β measured quality metrics
Reports measured quality of the student variants (perplexity, accuracy,
teacher divergence) β every number computed, none invented.
```bash
$ tensor-roll --workdir ./tr-quick evaluate
```
### `benchmark` β measured comparison table
Side-by-side measured comparison across variants (parameters, artifact
size, load time, latency, tokens/sec, perplexity). Backends that cannot
execute on the machine are reported as not executed, never estimated.
```bash
$ tensor-roll --workdir ./tr-quick benchmark
```
### `chat` β terminal inference REPL
Interactive inference against any exported variant, with optional
teacher/student side-by-side comparison.
Flags: `--variant {fp32,fp16,bf16,int8,int4}`, `--compare`, `--max-tokens`
```bash
$ tensor-roll --workdir ./tr-quick chat --variant int8
$ tensor-roll --workdir ./tr-quick chat --variant int4 --compare --max-tokens 64
```
### `sandbox` β OS-level sandbox supervisor
Runs a command inside the sandbox or demonstrates the allow/deny policy.
See [Sandbox](#sandbox) below.
Flags: `--cmd CMD`, `--demo`
```bash
$ tensor-roll sandbox --cmd "echo hello-from-the-jail"
$ tensor-roll sandbox --demo
jail: /tmp/trdoc/sandbox/jail netns_available=True
[ALLOW] ALLOW echo: policy pass
out: hello-tensor-roll
[ALLOW] ALLOW jail write+read: policy pass
out: jail-write-ok
[DENY] DENY shadow: absolute path outside jail rejected: /etc/shadow
[DENY] DENY curl: denied binary: curl
[DENY] DENY rm: denied binary: rm
audit log: /tmp/trdoc/sandbox/audit.log
```
### `demo` β full executable demonstration
Runs the entire pipeline end to end and drops into the inference REPL.
`--fast` uses reduced training steps (still real training).
Flags: `--fast`
```bash
$ tensor-roll demo --fast
```
### `doctor` β backend capability report
Prints the backend matrix: what is available, what executed, and what did
not β with reasons. Also writes `doctor_report.json` with `--json`.
```bash
$ tensor-roll doctor
$ tensor-roll doctor --json
```
---
## The `.trq` Container Format
Tensor Roll artifacts ship as `.trq` files β a real binary codec
(`tensor_roll/trq.py`), little-endian, versioned, fully checksummed:

```
magic 4 bytes b'TRQ1'
header_len u32
header JSON {version, arch, n_tensors, provenance, roll_meta, created}
per tensor:
name_len u16
name bytes
ndim u8
shape ndim Γ u64
dtype_code u8 0=fp32 1=fp16 2=bf16 3=int8 4=int4-packed
quant_code u8 0=none 1=int8-asym 2=int4-group32 3=fp16-cast 4=bf16-cast
meta_len u32
meta JSON {quantization metadata: scales/zeros (base64 f64),
roll metadata, codebooks}
data_len u64
data bytes
checksum 32 bytes sha256(data)
footer 32 bytes sha256(all preceding bytes)
```
`read_trq()` verifies the file checksum and every per-tensor checksum and
raises on any mismatch β corrupt files never load silently.
Inspect any artifact:
```bash
$ tensor-roll --workdir ./tr-quick inspect --what trq --trq student-int4.trq
```
## Sandbox

Inference and training tools execute inside a constrained OS environment β
the model never implicitly inherits unrestricted host permissions:
- **Process isolation** β dedicated process groups, killed on timeout
- **Filesystem jail** β absolute paths outside the jail are rejected
- **Resource limits** β rlimits on CPU, memory, and process count
- **Network policy** β network namespace isolation (`CLONE_NEWNET`)
- **Binary policy** β explicit allow/deny list per command
- **Environment filtering** β the sandbox does not inherit the host env
- **Audit logging** β every decision appended to a JSONL audit log
Terminal commands requested by the model pass through the sandbox
supervisor:
```
USER β TERMINAL β TENSOR ROLL β MODEL β TOOL REQUEST
β SANDBOX SUPERVISOR β POLICY β DENY | EXECUTE (isolated process)
```
## Measured Results
Every number below was measured by running the test suite and the pipeline
on this machine (2 CPUs, 7 GB RAM, no GPU) β nothing estimated:
| Check | Result |
|---|---|
| Test suite | **70 / 70 pass** (`python -m unittest discover -s tests`) |
| CLI commands | **14 / 14 live** (+ `doctor` backend report) |
| Validation compression | teacher **41.1K** β student **11.4K** = **27.8%** (target ratio 8.5/30 = **28.3%**) |
| Structural planning | virtual 30.19B-param manifest β **8.000B** plan, < 0.1 s, 35 MB peak RSS |
| `.trq` round-trip | all variants verified, checksums enforced |
| Sandbox | allow demonstrated, deny demonstrated, audit log written |
## Measured benchmarks
Every chart below plots numbers measured by running the pipeline and the
test suite on this machine (2 CPUs, 7 GB RAM, no GPU) β nothing estimated:

*Validation compression: teacher 41.1K β student 11.4K params (27.8%),
against the 28.3% target ratio (8.5B / 30B).*

*Measured student perplexity: INT8 8.39 vs INT4 15.13 β presented as recorded.*

*Virtual 30B structural test: 30.19B-param manifest β 8.000B plan in < 0.1 s,
35 MB peak RSS, zero weights allocated.*

*70/70 tests pass, 14/14 CLI commands live, 5/5 `.trq` variants round-trip
verified, crash-resume recovery byte-identical.*
## Layout
```
tensor_roll/ Python package: core, qsim, model, search, distill, quant,
trq, sandbox, boundary, cli, ooc, planner, fidelity,
doctor, ir, conformance
cudaq/ real CUDA-Q kernels (tensor_roll_kernels.py)
qsharp/ real Q# operations (TensorRoll.qs)
cuda/ CUDA GEMM sources
docs/ ARCHITECTURE.md, BOUNDARIES.md, demo assets
docs/img/ v1.0 product imagery
experiments/ measured experiment reports (JSON)
tests/ 70 tests, all green
```
## Author
**Ahmad Parr** β <Ahmedparr93@gmail.com>
## Version
**v1.0** (2026-09-29)
## License
AGPL-3.0-or-later β see [LICENSE](LICENSE). Every source file carries the
SPDX header `AGPL-3.0-or-later`, Copyright (C) 2026 SnapKitty Collective.
### πΌ Commercial License
SnapKitty code is free and open under **AGPL-3.0** for open-source use. Building a commercial product or service? A **proprietary commercial license** from Snapkitty Collective LLC lets you ship this code without the AGPL's source-sharing and network-use obligations.
**[β Get a commercial license](mailto:A.parr@belespritdaccord.uk?subject=Commercial%20license:%20tensor-roll)** Β· A.parr@belespritdaccord.uk
|