TwIL-LM3-Pro β€” LiteRT-LM

webAI-Official/TwIL-LM3-Pro converted to Google's LiteRT-LM (.litertlm) format for on-device inference. Requires LiteRT-LM 0.16 or newer.

TwIL-LM3-Pro is a 3.66B-parameter formal-logic reasoning model based on IBM Granite 4.2-3B. These bundles preserve its ChatML prompt format, pre-open the <think> block, and declare a thought channel so LiteRT-LM can expose or budget reasoning separately from the final answer.

Files

File Quantization Size Intended use
TwIL-LM3-Pro_int4.litertlm INT4 blockwise-32 + OCTAV linears; INT8 embedding 2.19 GB Mobile / smallest deployable build
TwIL-LM3-Pro_int8.litertlm Dynamic per-channel INT8 linears and embedding 3.76 GB Desktop / quality-oriented build

Both bundles use a 4,096-token KV cache and six prefill signatures: 1,024, 256, 64, 16, 4, and 1 tokens. The input embedding is externalized into its own bundle section to keep the primary graph below mobile mmap limits.

Usage

Install a recent LiteRT-LM runtime:

pip install "litert-lm>=0.16"

Run locally:

litert-lm run ./TwIL-LM3-Pro_int4.litertlm \
  --backend gpu \
  --thinking true \
  --prompt "Does 'All dogs are mammals. Rex is a dog.' entail 'Rex is a mammal'?"

Run directly from Hugging Face:

litert-lm run \
  --from-huggingface-repo webAI-Official/TwIL-LM3-Pro-LiteRT-LM \
  TwIL-LM3-Pro_int4.litertlm \
  --backend gpu \
  --thinking true \
  --prompt "What is the capital of France?"

For reasoning-heavy prompts, allow at least 2,048 output tokens. A generation that reaches its limit inside <think> may never produce a final answer. Use --thinking false for lower latency when reasoning is not needed.

Validation performed

The final repacked artifacts were checked with LiteRT-LM 0.16.1 on Linux CPU:

  • INT4, thinking disabled: What is the capital of France? β†’ Paris
  • INT8, thinking disabled: What is the capital of France? β†’ Paris
  • INT4, thinking budget 64: the runtime emitted a separate [thought] channel and answered the syllogism prompt with yes
  • Both metadata repacks proved all non-metadata sections byte-identical to the converter output

This is a conversion smoke test, not a benchmark. The bf16 Track A and Track B scores from the source model card must not be attributed to either quantized bundle until those evaluations are rerun through LiteRT-LM. In particular, reasoning models can be more sensitive to INT4 quantization than ordinary chat models; prefer INT8 when memory permits.

Conversion

Converted from the released model.safetensors, tokenizer, and configuration in webAI-Official/TwIL-LM3-Pro using:

  • litert-torch 0.9.3
  • litert-converter 0.4.0
  • ai-edge-quantizer 0.9.0
  • litert-lm-builder 0.16.1
  • transformers 5.14.1
  • john-rocky/hf-to-litertlm revision cb2ed3c53bd397acb82f132ee3d4b4313827ef30

The release uses the LiteRT-compatible Granite 4.2 Jinja template from that conversion project. It supports normal thinking on/off behavior and multi-turn history. The source checkpoint's extra reasoning_effort="low" template option is not exposed by this LiteRT bundle.

See CONVERSION.md and litertlm_manifest.json for the exact recipe and artifact hashes.

Limitations

  • Context is capped at 4,096 tokens in these on-device bundles, not the source checkpoint's inherited 131,072-token maximum.
  • Track A generations are often long. Mobile sessions need a sufficiently large output budget and enough memory for the selected backend.
  • Device-specific CPU/GPU performance and full delegation have not yet been measured for these TwIL weights.
  • The INT4 and INT8 quantized models have not yet been evaluated on the published formal-logic or held-out benchmark suites.

License

Released under the webAI Non-Commercial License ver. 1.0, matching the source TwIL-LM3-Pro model. The Granite 4.2 base attribution and Apache 2.0 license text are retained in apache-2.0-LICENSE.txt.

Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for webAI-Official/TwIL-LM3-Pro-LiteRT-LM