BarunAction-35M / source /docs /int8-quantization.md
harrrshall's picture
Release BarunAction-35M candidate-v2
5a46e5d verified
|
Raw
History Blame Contribute Delete
8.19 kB

BarunAction-35M CPU dynamic-int8 artifact

The retained int8 artifact is a version-bound CPU derivative of the exact candidate-v2 float checkpoint. It reduces the four-file core checkpoint from 141,440,943 to 58,619,256 bytes (58.56%) and passed a frozen 756-row accuracy-retention check. It did not demonstrate a speedup and is not portable outside its recorded runtime contract.

What is quantized

Format barun-cpu-dynamic-int8-v1 converts eligible torch.nn.Linear weights to per-tensor symmetric qint8 using an explicitly selected PyTorch CPU quantization engine. Activations are quantized dynamically inside each Linear operation.

For candidate-v2:

  • 87 internal Linear modules use qint8 weights;
  • lm_head remains FP32 and tied to embedding.weight;
  • token embeddings, RMS norms, RoPE buffers, attention math, residual operations, and KV state remain floating point; and
  • the serialized artifact retains the exact BarunLM architecture and tokenizer.

It is therefore a mixed-precision dynamic-int8 checkpoint, not an all-int8 model.

Runtime contract

The released artifact is accepted only under:

Field Required value
Operating system Darwin
CPU architecture ARM64
PyTorch 2.13.0
Quantization engine QNNPACK
Device CPU

The loader fails closed on a different platform, PyTorch version, engine, or CUDA request. Packed quantized state is not assumed to be portable across PyTorch releases or backends.

Artifact identity

File Bytes SHA-256
barun_config.json 448 9b3a1d71baa95a198744d250f9629231738d942570b8685c44307fd83dd33565
model.int8.pt 57,478,879 18102649bb4f8507ee754ae0c298e20064580833f05ca515faf5e04eac6f4488
tokenizer.json 1,136,031 70ded9605fccd09c2340ca7e225361eab0ae8b4dbbb0d6e26343ab5183979db6
quantization_manifest.json 3,898 f45c391d18d78758b0d62eeb562d139d24cfe0be3c8a409e3b40c25945d95c6b

The four core files total 58,619,256 bytes. The immutable W&B package also includes LICENSE, NOTICE, and MODEL_CARD.md, for 58,646,018 bytes total. The W&B identity is:

harshalsingh1223-gladium-ai/barunaction-35m/
barunaction-35m-candidate-v2-int8-darwin-arm64-qnnpack:v0

Its W&B digest is 8c3d8d30afc140e8c63461668ecb91a4.

Manifest and loader safeguards

The quantization manifest binds:

  • the three source float checkpoint hashes;
  • all int8 payload hashes and sizes;
  • the quantization algorithm and engine;
  • exact qint8 and FP32 Linear module inventories;
  • parameter accounting;
  • PyTorch, operating-system, and CPU-architecture versions; and
  • the embedding/output-head tying invariant.

Callers must supply the manifest SHA-256 out of band. An adjacent unpinned manifest is not accepted as authenticity evidence. Before torch.load, the verifier rejects missing/extra files, symlinks, oversized artifacts, hash mismatches, unsupported runtime fields, and unexpected state structure. It loads with weights_only=True, reconstructs the expected quantized graph, strict-loads the state, and verifies every intended qint8 weight and the retained FP32 tie.

weights_only=True narrows pickle execution risk; it is not a general defense against denial of service or lower-level deserializer vulnerabilities. Load only the hash-pinned artifact from a trusted channel.

Retrieve and verify

uv run --with 'wandb==0.28.1' wandb artifact get \
  --root ./models/barunaction-int8 \
  --type model \
  harshalsingh1223-gladium-ai/barunaction-35m/barunaction-35m-candidate-v2-int8-darwin-arm64-qnnpack:v0

barunaction verify-int8 \
  --checkpoint ./models/barunaction-int8 \
  --manifest-sha256 f45c391d18d78758b0d62eeb562d139d24cfe0be3c8a409e3b40c25945d95c6b

Inference requires the format and manifest identity explicitly:

barunaction infer \
  --checkpoint-format int8 \
  --checkpoint ./models/barunaction-int8 \
  --int8-manifest-sha256 f45c391d18d78758b0d62eeb562d139d24cfe0be3c8a409e3b40c25945d95c6b \
  --tools examples/barunaction_tools.example.json \
  --context examples/barunaction_empty_context.example.json \
  --now 2026-08-03T20:00:00+05:30 \
  --request "Turn on the flashlight" \
  --device cpu

There is no automatic float fallback.

Export a new compatible artifact

The implementation is generic over a compatible, hash-pinned BarunLM checkpoint:

barunaction export-int8 \
  --source-checkpoint ./models/BarunAction-35M \
  --source-hashes configs/barunaction/candidate-v2.json \
  --output ./models/barunaction-int8-new \
  --qengine qnnpack

The destination must not exist. The exporter prints a new manifest SHA-256; retain that digest outside the artifact directory. A newly exported artifact is not the released int8 checkpoint unless every payload and manifest hash matches this document.

Exact-output smoke

examples/barunaction_int8_smoke.example.json contains two simple expected Action IR cases. The smoke command runs the production compiler for both float and int8, requires both canonical ASTs to match the expected result, and reports raw output equality separately:

barunaction smoke-int8 \
  --source-checkpoint ./models/BarunAction-35M \
  --source-hashes configs/barunaction/candidate-v2.json \
  --int8-checkpoint ./models/barunaction-int8 \
  --manifest-sha256 f45c391d18d78758b0d62eeb562d139d24cfe0be3c8a409e3b40c25945d95c6b \
  --cases examples/barunaction_int8_smoke.example.json \
  --report ./int8-smoke-report.json

The released smoke passed 2/2 expected ASTs, and both float/int8 raw outputs matched. Its report SHA-256 is 1007b8ea9c5db56f1298893729a12e065fe21d91f232187864fddcd3882797ab.

Accuracy-retention result

The product-facing retention record is configs/barunaction/candidate-v2-int8.json, SHA-256 1228636074d7790b223eca1b25dd23ecfe0781b98f0e40df080b0608bd9f6201. The frozen protocol behind that record bound the exact source and int8 hashes, runtime, 756 development IDs, deterministic decoding, 192-token limit, minimum exact score 587/756, and maximum loss of 15 previously correct float rows before int8 scoring.

Result Value
Int8 AST exact 607/756 (80.29%)
Int8 parse valid 756/756
Int8 schema valid 755/756
Missing / generation failure / truncation 0 / 0 / 0
Versus frozen 602 float reference 8 fixed, 3 regressed, net +5
Same-host FP32 control 603/756
Int8 versus same-host FP32 7 fixed, 3 regressed, net +4

The sample-level transition file is benchmarks/evidence/int8-paired-outcomes.jsonl, SHA-256 e63e48dc8149b615ab03a14fb9fff9b1435ab59c6623e72c16b44d34d7c651ad.

Additional immutable receipt hashes:

Evidence SHA-256
Int8 aggregate 481a1e75210fa664cdcd0a55c0cd812393f79f1b74110dd9756f546701c4ef44
Int8 sample scores 80b4972efe3ea22d2f78887c7c671e6ee2845c642ec3d6b73fecad7c5849a337
Int8 raw predictions ad6ea7d41762ccdc34d018ef1158642c6bcada2c2c0682e356bebed9180fb4fc
Paired gate result 7e74efc58a52e16fc74e9c48eb64d86c1ce5d0bfd6459a7f1fd8ba66218fd688
Same-host FP32 comparison 83eb9e3967aee973cb1ed17e0602ed2e88838a836141ddb5f9ecd22c87a401e1

The int8 check was post-selection and used the same development population. Its higher point score does not establish that quantization generally improves accuracy.

Performance boundary

The two-case smoke measured 0.342 seconds total float inference and 0.590 seconds total int8 inference; int8 was slower. Order was not randomized, cases were not repeated, and peak memory, energy, cold/warm distributions, and a fixed production device were not measured. These numbers are diagnostic only, not an on-device latency benchmark.

PyTorch 2.13 reports the eager torch.ao.quantization.quantize_dynamic path as deprecated in favor of torchao. Deployment on another OS, architecture, PyTorch version, or engine requires a newly exported and fully re-evaluated artifact or a separately specified backend-neutral format.