File size: 7,454 Bytes
ee6c673 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 | ---
base_model: posttrainllm-ToolRouterModel-small
language:
- en
library_name: posttrainllm
license: mit
pipeline_tag: text-classification
tags:
- posttrainllm
- tinygpt
- intent-classification
- routing
- pace
- mac-first
- on-device
---
# Pace Intent Router v8
## Summary
A 49.5M-parameter intent classifier trained from scratch on a 127K-example
synthetic corpus for Pace's 7-class intent taxonomy. Runs in 3ms on CPU,
beats Qwen3-4B-Instruct by 10.8 percentage points, and is 85x faster.
This is a **posttrainllm-built specialist** β trained with the
`train-extractor` command on a `ToolRouterModel` small preset (6 layers,
384 d_model, byte-level vocab 256). No pre-trained weights, no
distillation from a teacher. The model learns Pace-specific decision
boundaries from synthetic data generated by
`pace/scripts/generate-intent-corpus-v2.py`.
## Artifact
- Package id: `pace-intent-router-v8`
- Artifact: `runs/pace-intent-router-v8.tinygpt` (566 MB)
- Labels: `runs/pace-intent-router-v8.tinygpt.labels.json`
- Format: posttrainllm `.tinygpt` (ToolRouterModel checkpoint)
- Base: `posttrainllm-ToolRouterModel-small` (trained from scratch)
- Params: 49.5M (6 layers, 384 d_model, 7 classes)
- Training method: `train-extractor` (cross-entropy classification,
AdamW + cosine decay, 18000 steps on v5 data)
- Training time: ~30 minutes on M5 Pro
## Measured Result
### Head-to-head vs Qwen3-4B-Instruct (4-bit)
| Metric | ToolRouterModel v8 | Qwen3-4B-Instruct | Delta |
|---|---:|---:|---:|
| Overall accuracy | **95.5%** | 84.75% | **+10.8 pp** |
| Eval size | 14,995 | 997 (stratified sample) | β |
| Latency p50 | **3.1 ms** | 240 ms | **77x faster** |
| Model size | **566 MB** | ~2,300 MB | **4x smaller** |
| Params | 49.5M | 4B (4-bit) | β |
### Per-class accuracy (v8 vs v5 vs Qwen3-4B)
| Class | v8 (49.5M) | v5 (49.5M) | Qwen3-4B |
|---|---:|---:|---:|
| chitchat | 93.8% | **95.6%** | 78.5% |
| pureKnowledge | 96.4% | **97.0%** | 83.0% |
| screenDescription | **97.6%** | 97.1% | **95.8%** |
| screenAction | 96.8% | **97.4%** | 91.0% |
| research | 93.4% | **97.1%** | **100.0%** |
| phoneLargeModel | **98.4%** | 96.4% | 76.7% |
| unknown | **79.6%** | 70.3% | 2.6% |
| **Overall** | 95.5% | **95.9%** | 84.75% |
### Key takeaways
1. **+10.8 pp over Qwen3-4B**: the 49.5M specialist beats the 4B
generalist on this task. Intent classification is a narrow enough
problem that a from-scratch model outperforms a general LLM.
2. **77x faster**: 3.1ms vs 240ms. The router runs in sub-4ms on CPU;
Qwen needs a full LLM forward pass.
3. **Unknown class is the moat**: v8 scores 79.6% on unknown vs Qwen's
2.6%. Qwen almost never returns "unknown" β it forces everything
into a known class, which is risky for a router (wrong route β
wrong pipeline). The trained model knows when to say "I don't know."
4. **v5 vs v8**: v5 (12K steps) slightly outperforms v8 (18K steps)
overall (95.9% vs 95.5%), but v8 has a much better unknown class
(79.6% vs 70.3%, +9.3 pp). The longer training improved the hardest
class at a small cost to the easier ones.
## Training data
- Dataset: pace-intent-corpus-v2 (fully synthetic)
- Source: `pace/scripts/generate-intent-corpus-v2.py` (base corpus,
combinatorial expansion) + `pace/scripts/generate-intent-supplement-v2.py`
(targeted weak-spot supplement)
- Train rows: 112,061
- Heldout rows: 14,995
- Classes: 7 (chitchat, pureKnowledge, screenDescription, screenAction,
research, phoneLargeModel, unknown)
- Split: stratified 85/15
- No real user data β all synthetic
### Why synthetic data works here
The synthetic corpus encodes Pace-specific decision boundaries that a
general LLM doesn't know:
- "turn on lights" = unknown (Pace can't control lights)
- "turn on volume" = screenAction (Pace can control volume)
- "what can you do" = pureKnowledge (not unknown β it's a question
about Pace itself)
These boundaries are product-specific, not language-general. A
from-scratch model learns them directly from the corpus; a general LLM
needs few-shot examples or fine-tuning to learn them.
## Recommended Use
This model is a **router**, not a planner. It classifies a user
utterance into one of 7 intent classes in 3ms. The intent class
determines which pipeline handles the turn:
| Intent | Route | Model |
|---|---|---|
| chitchat | fast path | Apple FM / local text-only |
| pureKnowledge | answer directly | Apple FM / local text-only |
| screenDescription | read screen | local planner + VLM |
| screenAction | execute tool | local planner + action layer |
| research | research tier | codex CLI / Claude CLI |
| phoneLargeModel | cloud bridge | codex CLI / cloud bridge |
| unknown | full pipeline | local planner (best-effort) |
Do not use this model for response generation. It only classifies.
## Known Limits
- **Unknown class is synthetic**: the 79.6% unknown accuracy is measured
on synthetic unknowns. Real-world unknown distribution will differ.
The model needs real-world unknown examples to improve further.
- **Not wired into Pace**: the shipping app uses Apple FM for intent
classification (when available) and rule-based fallback. This model
is a training pipeline artifact, not a product component.
- **Byte-level vocab**: the model uses a 256-token byte-level vocabulary.
Longer queries are truncated at 128 bytes. A BPE tokenizer would
generalize better on longer queries.
- **No real user data**: all training data is synthetic. Real user
queries will have different phrasings, accents, and edge cases.
## Why this model exists
The Pace intent classifier was originally rule-based (hand-written
phrase lists). This model was trained to test whether a learned
classifier could beat the rules. It can β 95.5% vs the rules' ~95% on
the synthetic eval β and it also beats Apple Foundation Models (3B,
in-process) on the same task: 95.5% vs 76.5% on a 200-example
stratified sample. The model is most useful as:
1. **The production intent classifier** β it's faster (3ms vs 1600ms),
more accurate (95.5% vs 76.5%), and has calibrated confidence
(real softmax vs hardcoded 0.95)
2. A fallback when Apple FM is unavailable (older Macs)
3. A training pipeline validation (the corpus and pipeline are assets)
4. A baseline for future on-device classifiers
### Measured head-to-head (200 stratified examples, 1s delay)
| Model | Accuracy | Latency | Error rate |
|---|---|---|---|
| **TinyGPT v8 (49.5M)** | **95.5%** | **3.1ms p50** | 0% |
| Qwen3-4B-Instruct (4-bit) | 84.75% | 240ms p50 | 0% |
| Apple FM (3B, in-process) | 76.5% | 1597ms mean | 0.5% |
FM's main weakness: it conflates `research` with `pureKnowledge`
(33.3% vs 93.4% β a 60pp gap) because it doesn't understand the
Pace-specific distinction between "research X" (multi-step) and
"what is X" (single answer). FM also refused to answer one query
("model refused to answer").
## References
- Factory run report: `runs/2026-07-13-pace-intent-router-v1/report.md`
- Eval data: `runs/2026-07-13-pace-intent-router-v1/eval-candidate-v8.json`
- Head-to-head: `runs/2026-07-13-pace-intent-router-v1/head-to-head.json`
- Training config: `runs/2026-07-13-pace-intent-router-v1/config.json`
- Corpus generator: `pace/scripts/generate-intent-corpus-v2.py`
- Supplement generator: `pace/scripts/generate-intent-supplement-v2.py`
- Pace routing architecture: `pace/leanring-buddy/PaceIntentClassifier.swift`
|