| --- |
| base_model: posttrainllm-ToolRouterModel-small |
| language: |
| - en |
| library_name: posttrainllm |
| license: mit |
| pipeline_tag: text-classification |
| tags: |
| - posttrainllm |
| - tinygpt |
| - intent-classification |
| - routing |
| - pace |
| - mac-first |
| - on-device |
| --- |
| |
| # Pace Intent Router v8 |
|
|
| ## Summary |
|
|
| A 49.5M-parameter intent classifier trained from scratch on a 127K-example |
| synthetic corpus for Pace's 7-class intent taxonomy. Runs in 3ms on CPU, |
| beats Qwen3-4B-Instruct by 10.8 percentage points, and is 85x faster. |
|
|
| This is a **posttrainllm-built specialist** β trained with the |
| `train-extractor` command on a `ToolRouterModel` small preset (6 layers, |
| 384 d_model, byte-level vocab 256). No pre-trained weights, no |
| distillation from a teacher. The model learns Pace-specific decision |
| boundaries from synthetic data generated by |
| `pace/scripts/generate-intent-corpus-v2.py`. |
| |
| ## Artifact |
| |
| - Package id: `pace-intent-router-v8` |
| - Artifact: `runs/pace-intent-router-v8.tinygpt` (566 MB) |
| - Labels: `runs/pace-intent-router-v8.tinygpt.labels.json` |
| - Format: posttrainllm `.tinygpt` (ToolRouterModel checkpoint) |
| - Base: `posttrainllm-ToolRouterModel-small` (trained from scratch) |
| - Params: 49.5M (6 layers, 384 d_model, 7 classes) |
| - Training method: `train-extractor` (cross-entropy classification, |
| AdamW + cosine decay, 18000 steps on v5 data) |
| - Training time: ~30 minutes on M5 Pro |
|
|
| ## Measured Result |
|
|
| ### Head-to-head vs Qwen3-4B-Instruct (4-bit) |
|
|
| | Metric | ToolRouterModel v8 | Qwen3-4B-Instruct | Delta | |
| |---|---:|---:|---:| |
| | Overall accuracy | **95.5%** | 84.75% | **+10.8 pp** | |
| | Eval size | 14,995 | 997 (stratified sample) | β | |
| | Latency p50 | **3.1 ms** | 240 ms | **77x faster** | |
| | Model size | **566 MB** | ~2,300 MB | **4x smaller** | |
| | Params | 49.5M | 4B (4-bit) | β | |
|
|
| ### Per-class accuracy (v8 vs v5 vs Qwen3-4B) |
|
|
| | Class | v8 (49.5M) | v5 (49.5M) | Qwen3-4B | |
| |---|---:|---:|---:| |
| | chitchat | 93.8% | **95.6%** | 78.5% | |
| | pureKnowledge | 96.4% | **97.0%** | 83.0% | |
| | screenDescription | **97.6%** | 97.1% | **95.8%** | |
| | screenAction | 96.8% | **97.4%** | 91.0% | |
| | research | 93.4% | **97.1%** | **100.0%** | |
| | phoneLargeModel | **98.4%** | 96.4% | 76.7% | |
| | unknown | **79.6%** | 70.3% | 2.6% | |
| | **Overall** | 95.5% | **95.9%** | 84.75% | |
|
|
| ### Key takeaways |
|
|
| 1. **+10.8 pp over Qwen3-4B**: the 49.5M specialist beats the 4B |
| generalist on this task. Intent classification is a narrow enough |
| problem that a from-scratch model outperforms a general LLM. |
| 2. **77x faster**: 3.1ms vs 240ms. The router runs in sub-4ms on CPU; |
| Qwen needs a full LLM forward pass. |
| 3. **Unknown class is the moat**: v8 scores 79.6% on unknown vs Qwen's |
| 2.6%. Qwen almost never returns "unknown" β it forces everything |
| into a known class, which is risky for a router (wrong route β |
| wrong pipeline). The trained model knows when to say "I don't know." |
| 4. **v5 vs v8**: v5 (12K steps) slightly outperforms v8 (18K steps) |
| overall (95.9% vs 95.5%), but v8 has a much better unknown class |
| (79.6% vs 70.3%, +9.3 pp). The longer training improved the hardest |
| class at a small cost to the easier ones. |
|
|
| ## Training data |
|
|
| - Dataset: pace-intent-corpus-v2 (fully synthetic) |
| - Source: `pace/scripts/generate-intent-corpus-v2.py` (base corpus, |
| combinatorial expansion) + `pace/scripts/generate-intent-supplement-v2.py` |
| (targeted weak-spot supplement) |
| - Train rows: 112,061 |
| - Heldout rows: 14,995 |
| - Classes: 7 (chitchat, pureKnowledge, screenDescription, screenAction, |
| research, phoneLargeModel, unknown) |
| - Split: stratified 85/15 |
| - No real user data β all synthetic |
|
|
| ### Why synthetic data works here |
|
|
| The synthetic corpus encodes Pace-specific decision boundaries that a |
| general LLM doesn't know: |
| - "turn on lights" = unknown (Pace can't control lights) |
| - "turn on volume" = screenAction (Pace can control volume) |
| - "what can you do" = pureKnowledge (not unknown β it's a question |
| about Pace itself) |
|
|
| These boundaries are product-specific, not language-general. A |
| from-scratch model learns them directly from the corpus; a general LLM |
| needs few-shot examples or fine-tuning to learn them. |
|
|
| ## Recommended Use |
|
|
| This model is a **router**, not a planner. It classifies a user |
| utterance into one of 7 intent classes in 3ms. The intent class |
| determines which pipeline handles the turn: |
|
|
| | Intent | Route | Model | |
| |---|---|---| |
| | chitchat | fast path | Apple FM / local text-only | |
| | pureKnowledge | answer directly | Apple FM / local text-only | |
| | screenDescription | read screen | local planner + VLM | |
| | screenAction | execute tool | local planner + action layer | |
| | research | research tier | codex CLI / Claude CLI | |
| | phoneLargeModel | cloud bridge | codex CLI / cloud bridge | |
| | unknown | full pipeline | local planner (best-effort) | |
|
|
| Do not use this model for response generation. It only classifies. |
|
|
| ## Known Limits |
|
|
| - **Unknown class is synthetic**: the 79.6% unknown accuracy is measured |
| on synthetic unknowns. Real-world unknown distribution will differ. |
| The model needs real-world unknown examples to improve further. |
| - **Not wired into Pace**: the shipping app uses Apple FM for intent |
| classification (when available) and rule-based fallback. This model |
| is a training pipeline artifact, not a product component. |
| - **Byte-level vocab**: the model uses a 256-token byte-level vocabulary. |
| Longer queries are truncated at 128 bytes. A BPE tokenizer would |
| generalize better on longer queries. |
| - **No real user data**: all training data is synthetic. Real user |
| queries will have different phrasings, accents, and edge cases. |
|
|
| ## Why this model exists |
|
|
| The Pace intent classifier was originally rule-based (hand-written |
| phrase lists). This model was trained to test whether a learned |
| classifier could beat the rules. It can β 95.5% vs the rules' ~95% on |
| the synthetic eval β and it also beats Apple Foundation Models (3B, |
| in-process) on the same task: 95.5% vs 76.5% on a 200-example |
| stratified sample. The model is most useful as: |
|
|
| 1. **The production intent classifier** β it's faster (3ms vs 1600ms), |
| more accurate (95.5% vs 76.5%), and has calibrated confidence |
| (real softmax vs hardcoded 0.95) |
| 2. A fallback when Apple FM is unavailable (older Macs) |
| 3. A training pipeline validation (the corpus and pipeline are assets) |
| 4. A baseline for future on-device classifiers |
|
|
| ### Measured head-to-head (200 stratified examples, 1s delay) |
|
|
| | Model | Accuracy | Latency | Error rate | |
| |---|---|---|---| |
| | **TinyGPT v8 (49.5M)** | **95.5%** | **3.1ms p50** | 0% | |
| | Qwen3-4B-Instruct (4-bit) | 84.75% | 240ms p50 | 0% | |
| | Apple FM (3B, in-process) | 76.5% | 1597ms mean | 0.5% | |
|
|
| FM's main weakness: it conflates `research` with `pureKnowledge` |
| (33.3% vs 93.4% β a 60pp gap) because it doesn't understand the |
| Pace-specific distinction between "research X" (multi-step) and |
| "what is X" (single answer). FM also refused to answer one query |
| ("model refused to answer"). |
|
|
| ## References |
|
|
| - Factory run report: `runs/2026-07-13-pace-intent-router-v1/report.md` |
| - Eval data: `runs/2026-07-13-pace-intent-router-v1/eval-candidate-v8.json` |
| - Head-to-head: `runs/2026-07-13-pace-intent-router-v1/head-to-head.json` |
| - Training config: `runs/2026-07-13-pace-intent-router-v1/config.json` |
| - Corpus generator: `pace/scripts/generate-intent-corpus-v2.py` |
| - Supplement generator: `pace/scripts/generate-intent-supplement-v2.py` |
| - Pace routing architecture: `pace/leanring-buddy/PaceIntentClassifier.swift` |
|
|