decider-0.8b — typed decisions in LiteRT-LM

To route a support ticket or score a yes/no question, supply a state and typed questions and read option probabilities. The example request asks who should handle a duplicate-charge ticket, whether a refund was requested, and how frustrated the customer is. The upstream fp32 response assigns billing probability 0.9996 and refund-request probability 0.9611; the fp16 reference produces the same rounded values for this example. The complete 40-fixture corpus and fp32 oracle outputs are included.

This is a third-party conversion of Mapika/decider-0.8b, fine-tuned from Qwen/Qwen3.5-0.8B-Base, for litert-community/decider-0.8b-LiteRT. Upstream makes one-pass typed decisions with calibrated option probabilities: choice (2–255 options), score (2–10 levels), and noul (probability of yes). It does not generate answer text. The request/answer shape is the decider package’s /v1/systemone endpoint, which follows TypeSafe Jev’s wire format. The CPU LiteRT reference implements these decisions with independent rows, exact-fit chunked prefill and one final decode per row; a score question expands into one yes/no row for each level.

The text-only architecture has 18 linear-attention and 6 full-attention layers. This export allows at most 4096 tokens per rendered row. Requires LiteRT-LM ≥0.15 for generalized hybrid-state binding; describe and engine operation were verified on 0.17.1. The reference graph reader uses ai-edge-litert 2.2.0. The fp16 file is recommended for desktop/CPU probability fidelity. The dynamic-int8 file runs on CPU and on the tested Mac Metal and Galaxy S26 CL GPUs; GPU decision readout requires the single-chunk rule below.

Files

File Weight recipe / intended use Bytes SHA-256
decider-0.8b_fp16.litertlm FP16 casting of fully connected and embedding weights; float compute. Recommended for upstream probability fidelity on desktop/CPU. 1,625,055,920 477c1dcadbb229833cbcb9613f3436ad7057dcdbc4de2038ca3e8f4ec0485f25
decider-0.8b_int8.litertlm Dynamic int8 fully connected weights + int8 CHANNELWISE embedding, the same recipe as litert-community/Qwen3.5-0.8B; Android CPU or GPU. 963,184,864 3e3c03da33034115dade6094c8ae94a90a7169d2a7cb4c6f8145e2dcc49a143b

Convolutions and the delta rule remain float in both files. Each bundle has an identity prompt template, no start token, stop token 248044, ExecutorMetadata for 48 state buffers, and an fp32 activation preference. File checksums are in SHA256SUMS. The published int8 filename identifies the dynamic-int8 build, not the earlier unpublished weight-only-int8 build.

Correctness and probability preservation

Forty synthetic English state/question fixtures produce 120 independent rows, including a 255-option choice. The oracle is pinned upstream Decider with torch fp32 on Apple M4 Max CPU, four threads. The graph gate uses identical piecewise token IDs, label IDs, T = 1.03 and upstream assembly. Δp is each row’s maximum absolute difference across its unrounded option probabilities; p95 and median summarize those row maxima. GPU rows use ai-edge-litert 2.2.0 CompiledModel Metal with fp32 activations enforced and full acceleration. These are conversion-agreement measurements, not a real-traffic calibration or task-accuracy evaluation.

File Backend / scheme Rows Argmax identical Ties Max |Δp| p95 |Δp| Median |Δp| Rows > 0.02
fp16 CPU / exact-fit 120 120/120 0 6.61611557e-06 3.06963921e-06 1.1920929e-07 0
int8 CPU / exact-fit 120 117/120 0 0.159999788 0.067937851 0.00210475922 27
int8 CPU / one padded prefill 119 116/119 0 0.129208088 0.0709148347 0.00181061029 25
int8 Metal GPU / one padded prefill 119 118/119 0 0.0731657743 0.0250394106 0.000544799026 10

The int8 file's probabilities differ from the fp32 model on CPU because its fully connected layers dynamically quantize activations: exact-fit max |Δp| is 0.160 and p95 is 0.068, with 27/120 rows above 0.02. The GPU computes with the int8 weights in float; on Mac Metal its padded max |Δp| is 0.073 and p95 is 0.025, with 10/119 rows above 0.02. Use fp16 when preserving the upstream probabilities matters. The CPU padded control is 116/119 argmax, max 0.129 and p95 0.071; it is not numerically identical to CPU exact-fit.

Every output is finite. No oracle top-two gap is within the 1e-4 tie threshold. For fp16, the reference CLI matches all 40 fixture answer dictionaries within 1e-4 after upstream rounding (36 exact), and all 40 usage counts match; this CLI statement does not apply to int8. The 119-row padded gates exclude G01/00, the 2844-token wide-choice row. Its historical int8 GPU multi-chunk error was 0.0238454, but it is not part of the padded GPU calibration; route such long rows to CPU. Android label-probability calibration was not measured: the S26 results below establish runtime operation, not equality to the Mac probabilities.

Reference usage

The included CPU reference implementation, copied unchanged from the fp16 release reference, accepts state and questions, and prints the upstream response shape: model, answers, usage. Model diagnostics go to stderr. Using Python 3.12 from the repository root, install the checked host dependencies and download only the pinned tokenizer and decision configuration:

python -B -m venv .venv
.venv/bin/python -B -m pip install ai-edge-litert==2.2.0 litert-lm==0.17.1 numpy tokenizers huggingface_hub
HF_HUB_DISABLE_XET=1 HF_HOME=.cache/hf .venv/bin/python -B - <<'PY'
from huggingface_hub import snapshot_download
snapshot_download('Mapika/decider-0.8b', revision='1ea54127d3bd52f6d753d9257b32a6380b873907',
                  allow_patterns=['tokenizer.json', 'decider_config.json'], local_dir='tokenizer', max_workers=2)
PY
.venv/bin/python -B reference/systemone_litert.py \
  --bundle decider-0.8b_fp16.litertlm --tokenizer tokenizer \
  --request reference/fixtures/systemone_example.json
  1. Render raw state/question text with the pinned decider functions: state-first layout, independent rows, isolated score levels.
  2. Add no BOS or role markers; render message content verbatim.
  3. The answer slot is each row’s last token: prefill ids[:-1], then decode ids[-1] once.
  4. Read only label_table(tokenizer)[1][:nopts] from the slot logits, preserving piecewise IDs for wide choices.
  5. Apply softmax over those labels at T = 1.03, then use upstream assembly and usage counting.

The graph reader zeroes state for each row and carries all 48 buffers across greedy exact-fit prefill chunks of 1024,512,256,128,64,32,16,8,4,2,1 tokens, using absolute positions and causal masks. Prefill returns state only; the final decode returns logits. Whole-string tokenization matched upstream IDs in 119/120 rows; the 255-option row differed, so preserve the reference’s piecewise construction. The bundle helper unpacks once to a repository-local .cache/readout directory and reuses it; remove that temporary cache after use when reclaiming disk space.

GPU feeding rule: zero all state for each independent row and feed ids[:-1] in ONE padded prefill chunk, choosing the smallest exported prefill signature with length >= len(ids)-1; then decode the final real token at position len(ids)-1. Put valid token IDs first and fill the remaining token slots with ID 0. Valid positions are 0..N-1, padded positions are 0, valid query masks are causal (0 for allowed entries and -1e30 otherwise), and padded query masks are entirely -1e30; the exported guard identifies valid prefill tokens with input_ids != 0. Keep fp32 GPU activations. The 119 measured rows contain no natural token ID 0; inputs containing that ID are outside the measured pad-guard contract.

Chaining several prefill chunks on the GPU corrupts the carried state on this hybrid graph. In the Mac Metal feeding experiment, maximum |Δp| versus CPU exact-fit was 0.177 with multi-chunk prefill, 9.8e-6 with one padded chunk, and 5.9e-6 with decode-walk; CPU results agreed across all three schemes within 8.9e-6. Rows longer than 1024 tokens therefore run on CPU, even though the exported state cache holds 4096 tokens. The included reference is CPU-only and preserves the exact-fit contract; GPU integration must implement the rule above. A successful native GPU generation benchmark does not verify its internal feeding plan or option-probability readout.

Prompt and assembly code is vendored from Mapika/decider commit c4daaac28af9fea95d627015cffa2dd5a5926ee6: prompt.py SHA-256 5a42134cf470c566e34ac38fb10e63851c21a4bb739475eef797849bcfe460a3; systemone.py SHA-256 627fc365583c92328b7b21d481b307eb4d2bfd27411d38cd6710ab3adf629a9a. The infer.py excerpt retains upstream request dataclasses and None handling. Their Apache-2.0 license is included.

Engine scoring: LiteRT-LM RunTextScoring is available through the C/Python API, one target per call. On this hybrid architecture use a fresh process/engine for each scoring call: scoring advances session state, and shared-engine prefix reuse can retain stale recurrent/conv state. The fp16 six-row hermetic control preserved all six argmaxes but had max Δp 0.004485; it does not establish exact probability parity. The shipped CPU readout is the graph path above; no equivalent hermetic engine-scoring probability claim is made for the dynamic-int8 file.

Performance

File Device / backend LiteRT-LM version Processed prefill / decode tokens Prefill tok/s TTFT s Decode tok/s¹ Reported init s Peak VmHWM MiB Conditions
fp16 Apple M4 Max / CPU 4 threads 0.17.1 256 / 8 533.23 0.5335 18.82 42.3812 not measured 3 measured + 1 excluded warmup; cache no; quiet
int8 Apple M4 Max / CPU 4 threads 0.17.1 256 / 8 695.41 0.3917 42.48 40.3517 not measured 3 measured + 1 excluded warmup; cache no; quiet
int8 Galaxy S26 SM-S942Q / GPU 0.16.0 S4 kit 262 / 3834 300.48 0.91 25.29 78.28923 4552.68359 1 cold run; caches removed; thermal-checked
int8 Galaxy S26 SM-S942Q / GPU 0.16.0 S4 kit 503 / 3 567.19 0.97 12.05 97.37262 5317.46484 1 cold run; caches removed; thermal-checked
int8 Galaxy S26 SM-S942Q / CPU 0.16.0 S4 kit 262 / 3834 287.75 0.94 36.89 13.81788 1941.13281 1 cold run; caches removed; thermal-checked
fp16 Galaxy S26 SM-S942Q / CPU 0.16.0 S4 kit 262 / 3834 44.94 5.92 10.98 18.08928 6093.80859 1 cold run; caches removed; thermal-checked

Mac measurements: 2026-09-23, macOS 27.0 (26A428), quiet machine (supervisor checked), with no other Codex run at launch. Prefill, TTFT and decode are the CLI means of three measured iterations after one excluded warmup, with CPU 4 threads and --cache no. Mac init is the CLI-reported value from the first measured iteration, not a separately timed constructor. The fp16 quiet row is reused from the earlier measurement of this exact SHA; int8 was measured for this release.

S26: litert_lm_advanced_main v0.16.0 S4 kit (2026-08-17), CPU sampler, native default CPU thread count. Each row is one cold run, with caches deleted and SKIN <40°C, thermal status <=1, and uncapped cpu0/4/7 verified before launch. The int8 GPU delegated 188522/188522 ops across all 12 subgraphs on both prompts; both GPU runs and the CPU control exited 0. VmHWM is sampled in kB and divided by 1024 to report MiB, including initialization, prefill and the full decode. Source prompts contain 262 and 503 tokens; the table reports the runtime's processed counts. ¹ Decode is informational only: generated text has no decision meaning, and the int8 GPU503 run ended normally after 3 decode tokens. These are generation-runtime benchmarks, not System One request latency or an Android probability-parity gate.

At about 260 tokens the phone GPU is not materially faster than the CPU for this prefill-dominated model (300.48 versus 287.75 tok/s; TTFT 0.91 versus 0.94 s), while it costs 78–97 s of engine initialization and roughly 2–3× the memory (4553–5317 MiB versus 1941 MiB). Its benefit is calibration, as measured on Mac Metal above; S26 calibration itself is unmeasured. These single cold rows do not establish a throughput ranking between runtimes.

The fp16 S26 attempt outcomes were one run per prompt, without retry:

Prompt row Source tokens Cold run Status Named reason
E03/00 262 1 PASS Completed; 44.94 tok/s prefill, 5.92 s TTFT, 6094 MiB peak
F01/01 503 1 BLOCKED Device left ADB; the phone logged a framework restart (SYSTEM_RESTART 08:51:50 JST, kernel uptime unchanged) and returned within minutes; not retried

The fp16 503-token run lost ADB at 08:51:48 JST, 92.706 s after launch, with the process at 7740 MiB VmHWM in the last sample (a lower bound). The recorded event was a framework restart, not a kernel reboot. The exact fp16 file also failed GPU engine creation on this S26 CL kit because partial delegation was refused. The fp16 file is not recommended on phones; use it on desktop/CPU when upstream probability fidelity matters.

Not shipped

The previous weight-only int8 build (964,791,472 B, SHA-256 9363b8100afbd11690b54f4bafbfa0f053a391beaabee9a02923b4c254a59324) is a different file from this release's dynamic int8 and remains unpublished after two Galaxy S26 kernel panics, whose cause was not identified.

The fp16-FC + dynamic-int8-head + int8-embedding form (V7c) is also not published in this release: its Mac Metal padded p95 |Δp| was 0.005, and it fully delegated on S26 GPU, but the 262-token run peaked at 6175 MiB (6.47 GB decimal).

Limitations

This is a small English-language model without a reasoning stage. Procedural rules embedded in a question may be ignored; direct questions with described options are the tested form. Knowledge-heavy choices remain limited by the base model. Upstream calibration uses public datasets and teacher-labelled probes, not a user’s traffic, and training examples may carry the Qwen3.5-27B teacher’s biases. In retained fixture C04, the record says approval is false while the upstream model gives yes probability 0.5692; that is upstream behavior, not a conversion defect. See the upstream limitations.

The graph’s 4096-token row limit is lower than the upstream maximum state length. Schema-first decide(), multiple slots in a row and cached-prefix interfaces are outside the tested contract. Generated text from runtime benchmarks is not a model capability claim. Dynamic-int8 GPU runtime operation was tested on Mac Metal and one Galaxy S26; probability calibration was measured only on Mac Metal using one padded prefill chunk. GPU rows longer than 1024 tokens must be routed to CPU. Int8 CPU probabilities differ from the upstream fp32 oracle as quantified above.

License and changes

Apache-2.0 is inherited from both the linked Qwen base model and Mapika finetune; the vendored reference package is also Apache-2.0. Source checkpoint revision: 1ea54127d3bd52f6d753d9257b32a6380b873907. Base revision checked for license: dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68.

Changes: converted the text decoder from safetensors to LiteRT flatbuffers; for the fp16 file, cast fully connected and embedding weights to fp16 while retaining float compute; for the int8 file, apply the dynamic-int8 recipe to fully connected weights and int8 CHANNELWISE embedding while leaving convolutions and the delta rule float; repackaged the tokenizer unchanged; replaced the unused stock chat template with an identity template; added ExecutorMetadata and fp32 activation preference. No vision or MTP is included; the source text checkpoint contains no such tensors. Conversion uses the owner’s hybrid patch on litert-torch fork commit 115a13607c730c81018bb9789138a3e5e5119e3d. This community conversion is not an official Mapika or Qwen release.

Reproduction (converter, gates, fixtures): hf-to-litertlm decider_work/.

Downloads last month
222
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for litert-community/decider-0.8b-LiteRT

Finetuned
(2)
this model