--- base_model: posttrainllm-ToolRouterModel-small language: - en library_name: posttrainllm license: mit pipeline_tag: text-classification tags: - posttrainllm - tinygpt - intent-classification - routing - pace - mac-first - on-device --- # Pace Intent Router v8 ## Summary A 49.5M-parameter intent classifier trained from scratch on a 127K-example synthetic corpus for Pace's 7-class intent taxonomy. Runs in 3ms on CPU, beats Qwen3-4B-Instruct by 10.8 percentage points, and is 85x faster. This is a **posttrainllm-built specialist** — trained with the `train-extractor` command on a `ToolRouterModel` small preset (6 layers, 384 d_model, byte-level vocab 256). No pre-trained weights, no distillation from a teacher. The model learns Pace-specific decision boundaries from synthetic data generated by `pace/scripts/generate-intent-corpus-v2.py`. ## Artifact - Package id: `pace-intent-router-v8` - Artifact: `runs/pace-intent-router-v8.tinygpt` (566 MB) - Labels: `runs/pace-intent-router-v8.tinygpt.labels.json` - Format: posttrainllm `.tinygpt` (ToolRouterModel checkpoint) - Base: `posttrainllm-ToolRouterModel-small` (trained from scratch) - Params: 49.5M (6 layers, 384 d_model, 7 classes) - Training method: `train-extractor` (cross-entropy classification, AdamW + cosine decay, 18000 steps on v5 data) - Training time: ~30 minutes on M5 Pro ## Measured Result ### Head-to-head vs Qwen3-4B-Instruct (4-bit) | Metric | ToolRouterModel v8 | Qwen3-4B-Instruct | Delta | |---|---:|---:|---:| | Overall accuracy | **95.5%** | 84.75% | **+10.8 pp** | | Eval size | 14,995 | 997 (stratified sample) | — | | Latency p50 | **3.1 ms** | 240 ms | **77x faster** | | Model size | **566 MB** | ~2,300 MB | **4x smaller** | | Params | 49.5M | 4B (4-bit) | — | ### Per-class accuracy (v8 vs v5 vs Qwen3-4B) | Class | v8 (49.5M) | v5 (49.5M) | Qwen3-4B | |---|---:|---:|---:| | chitchat | 93.8% | **95.6%** | 78.5% | | pureKnowledge | 96.4% | **97.0%** | 83.0% | | screenDescription | **97.6%** | 97.1% | **95.8%** | | screenAction | 96.8% | **97.4%** | 91.0% | | research | 93.4% | **97.1%** | **100.0%** | | phoneLargeModel | **98.4%** | 96.4% | 76.7% | | unknown | **79.6%** | 70.3% | 2.6% | | **Overall** | 95.5% | **95.9%** | 84.75% | ### Key takeaways 1. **+10.8 pp over Qwen3-4B**: the 49.5M specialist beats the 4B generalist on this task. Intent classification is a narrow enough problem that a from-scratch model outperforms a general LLM. 2. **77x faster**: 3.1ms vs 240ms. The router runs in sub-4ms on CPU; Qwen needs a full LLM forward pass. 3. **Unknown class is the moat**: v8 scores 79.6% on unknown vs Qwen's 2.6%. Qwen almost never returns "unknown" — it forces everything into a known class, which is risky for a router (wrong route → wrong pipeline). The trained model knows when to say "I don't know." 4. **v5 vs v8**: v5 (12K steps) slightly outperforms v8 (18K steps) overall (95.9% vs 95.5%), but v8 has a much better unknown class (79.6% vs 70.3%, +9.3 pp). The longer training improved the hardest class at a small cost to the easier ones. ## Training data - Dataset: pace-intent-corpus-v2 (fully synthetic) - Source: `pace/scripts/generate-intent-corpus-v2.py` (base corpus, combinatorial expansion) + `pace/scripts/generate-intent-supplement-v2.py` (targeted weak-spot supplement) - Train rows: 112,061 - Heldout rows: 14,995 - Classes: 7 (chitchat, pureKnowledge, screenDescription, screenAction, research, phoneLargeModel, unknown) - Split: stratified 85/15 - No real user data — all synthetic ### Why synthetic data works here The synthetic corpus encodes Pace-specific decision boundaries that a general LLM doesn't know: - "turn on lights" = unknown (Pace can't control lights) - "turn on volume" = screenAction (Pace can control volume) - "what can you do" = pureKnowledge (not unknown — it's a question about Pace itself) These boundaries are product-specific, not language-general. A from-scratch model learns them directly from the corpus; a general LLM needs few-shot examples or fine-tuning to learn them. ## Recommended Use This model is a **router**, not a planner. It classifies a user utterance into one of 7 intent classes in 3ms. The intent class determines which pipeline handles the turn: | Intent | Route | Model | |---|---|---| | chitchat | fast path | Apple FM / local text-only | | pureKnowledge | answer directly | Apple FM / local text-only | | screenDescription | read screen | local planner + VLM | | screenAction | execute tool | local planner + action layer | | research | research tier | codex CLI / Claude CLI | | phoneLargeModel | cloud bridge | codex CLI / cloud bridge | | unknown | full pipeline | local planner (best-effort) | Do not use this model for response generation. It only classifies. ## Known Limits - **Unknown class is synthetic**: the 79.6% unknown accuracy is measured on synthetic unknowns. Real-world unknown distribution will differ. The model needs real-world unknown examples to improve further. - **Not wired into Pace**: the shipping app uses Apple FM for intent classification (when available) and rule-based fallback. This model is a training pipeline artifact, not a product component. - **Byte-level vocab**: the model uses a 256-token byte-level vocabulary. Longer queries are truncated at 128 bytes. A BPE tokenizer would generalize better on longer queries. - **No real user data**: all training data is synthetic. Real user queries will have different phrasings, accents, and edge cases. ## Why this model exists The Pace intent classifier was originally rule-based (hand-written phrase lists). This model was trained to test whether a learned classifier could beat the rules. It can — 95.5% vs the rules' ~95% on the synthetic eval — and it also beats Apple Foundation Models (3B, in-process) on the same task: 95.5% vs 76.5% on a 200-example stratified sample. The model is most useful as: 1. **The production intent classifier** — it's faster (3ms vs 1600ms), more accurate (95.5% vs 76.5%), and has calibrated confidence (real softmax vs hardcoded 0.95) 2. A fallback when Apple FM is unavailable (older Macs) 3. A training pipeline validation (the corpus and pipeline are assets) 4. A baseline for future on-device classifiers ### Measured head-to-head (200 stratified examples, 1s delay) | Model | Accuracy | Latency | Error rate | |---|---|---|---| | **TinyGPT v8 (49.5M)** | **95.5%** | **3.1ms p50** | 0% | | Qwen3-4B-Instruct (4-bit) | 84.75% | 240ms p50 | 0% | | Apple FM (3B, in-process) | 76.5% | 1597ms mean | 0.5% | FM's main weakness: it conflates `research` with `pureKnowledge` (33.3% vs 93.4% — a 60pp gap) because it doesn't understand the Pace-specific distinction between "research X" (multi-step) and "what is X" (single answer). FM also refused to answer one query ("model refused to answer"). ## References - Factory run report: `runs/2026-07-13-pace-intent-router-v1/report.md` - Eval data: `runs/2026-07-13-pace-intent-router-v1/eval-candidate-v8.json` - Head-to-head: `runs/2026-07-13-pace-intent-router-v1/head-to-head.json` - Training config: `runs/2026-07-13-pace-intent-router-v1/config.json` - Corpus generator: `pace/scripts/generate-intent-corpus-v2.py` - Supplement generator: `pace/scripts/generate-intent-supplement-v2.py` - Pace routing architecture: `pace/leanring-buddy/PaceIntentClassifier.swift`