# SAM Benchmark v1 — Full Report ## Methodology This benchmark evaluates structured action generation: given a natural language task and a JSON schema, produce a syntactically valid JSON sequence of operations that matches a reference answer. **Test suite**: 200 prompts, deterministically generated from SEED=42. The exact prompts are saved to `test_suite.json` for reproducibility and audit. **Tier distribution** (current suite): - Tier 1 — Atomic — single op: **40 cases** (20.0%) - Tier 2 — Compound — 2-3 ops same domain: **40 cases** (20.0%) - Tier 3 — Noisy — typos / casing / fillers: **40 cases** (20.0%) - Tier 4 — Long chain — 4+ ops: **40 cases** (20.0%) - Tier 5 — Cross-domain workflow: **40 cases** (20.0%) **Domain distribution** (current suite): - `WORKFLOW`: 49 cases (24.5%) - `ROS`: 46 cases (23.0%) - `DB`: 27 cases (13.5%) - `HTTP`: 25 cases (12.5%) - `VEHICLE`: 17 cases (8.5%) - `MQTT`: 15 cases (7.5%) - `ECOMMERCE`: 8 cases (4.0%) - `HOME`: 7 cases (3.5%) - `FILE`: 4 cases (2.0%) - `CAL`: 2 cases (1.0%) **Metrics**: - `JSON Valid`: output parses as valid JSON ∈ [0,100]% - `Op Correct`: output is a list of dicts AND every op name matches the expected set ∈ [0,100]% - `Exact Match`: full structural match with reference, with float tolerance 0.02 for numerics ∈ [0,100]% - `Latency`: mean wall-clock time per prompt in milliseconds (single-thread, no batching) - `Size`: model file size on disk (MB) **Important disclosures**: 1. SAM is **trained on the same domain distribution as this benchmark** (multi-domain JSON ops). General-purpose LLMs are evaluated zero-shot. This is a deliberate comparison of "specialist 35M model" vs "generalist N-B model" on a structured task, NOT a claim of general superiority. 2. **Tier distribution is currently imbalanced** toward Tier 1 (atomic). Aggregate exact-match numbers are dominated by easy cases. Always read the per-tier breakdown below for fair comparison. 3. **Domain distribution is currently imbalanced** toward ROS (robotics). This reflects SAM's lineage from Foros (ROS-only). Per-domain breakdowns reveal each model's actual strength. 4. All LLMs are loaded from publicly available HuggingFace weights with no fine-tuning. Prompt template is identical across all LLMs. ## Headline results (aggregate across all 200 cases) | Rank | Model | Params | Size (MB) | JSON Valid | Op Correct | Exact Match | Latency (ms) | |---|---|---|---|---|---|---|---| | 1 | **[SAM]** SAM | 33.4M | 127.4 | 88.0% | 63.5% | 18.0% | 715.4 | | 2 | Qwen2.5-3B | 3085.9M | 5886.0 | 100.0% | 82.5% | 16.0% | 2061.1 | | 3 | DeepSeek-R1-Qwen3-8B-GGUF | 8000.0M | 4794.9 | 98.0% | 79.0% | 11.5% | 31737.9 | | 4 | TinyLlama-1.1B | 1100.0M | 2098.2 | 92.0% | 55.5% | 9.0% | 4199.3 | | 5 | Qwen2.5-1.5B | 1543.7M | 2944.4 | 100.0% | 78.5% | 9.0% | 1852.4 | | 6 | Qwen2.5-0.5B | 494.0M | 942.3 | 70.5% | 30.0% | 2.5% | 1937.2 | | 7 | SmolLM2-1.7B | 1711.4M | 3264.2 | 89.0% | 38.5% | 2.5% | 1694.9 | | 8 | SmolLM2-360M | 361.8M | 690.1 | 89.0% | 0.0% | 0.0% | 2042.9 | | 9 | Llama-3.2-1B | ? | 0.0 | 0.0% | 0.0% | 0.0% | 0.0 | | 10 | Gemma-2-2B | ? | 0.0 | 0.0% | 0.0% | 0.0% | 0.0 | | 11 | Llama-3.2-3B | ? | 0.0 | 0.0% | 0.0% | 0.0% | 0.0 | | 12 | Phi-3.5-mini-3.8B | ? | 0.0 | 0.0% | 0.0% | 0.0% | 0.0 | | 13 | Llama-3.2-3B-uns | ? | 0.0 | 0.0% | 0.0% | 0.0% | 0.0 | | 14 | Gemma-3-270M-GGUF | 270.0M | 241.4 | 92.0% | 2.0% | 0.0% | 6322.5 | | 15 | Ministral-3-8B-GGUF | ? | 0.0 | 0.0% | 0.0% | 0.0% | 0.0 | ## Per-tier exact-match breakdown (all models) _This is the breakdown that matters for fair comparison. Each model's exact-match score is shown for each difficulty tier._ | Model | T1 atomic | T2 compound | T3 noisy | T4 long | T5 cross | |---|---|---|---|---|---| | **[SAM]** SAM | 37.5% (n=40) | 27.5% (n=40) | 25.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | | Qwen2.5-3B | 40.0% (n=40) | 12.5% (n=40) | 27.5% (n=40) | 0.0% (n=40) | 0.0% (n=40) | | DeepSeek-R1-Qwen3-8B-GGUF | 27.5% (n=40) | 7.5% (n=40) | 22.5% (n=40) | 0.0% (n=40) | 0.0% (n=40) | | TinyLlama-1.1B | 27.5% (n=40) | 12.5% (n=40) | 5.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | | Qwen2.5-1.5B | 30.0% (n=40) | 12.5% (n=40) | 2.5% (n=40) | 0.0% (n=40) | 0.0% (n=40) | | Qwen2.5-0.5B | 12.5% (n=40) | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | | SmolLM2-1.7B | 12.5% (n=40) | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | | SmolLM2-360M | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | | Gemma-3-270M-GGUF | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | 0.0% (n=40) | ## Per-tier JSON-valid breakdown (all models) | Model | T1 atomic | T2 compound | T3 noisy | T4 long | T5 cross | |---|---|---|---|---|---| | **[SAM]** SAM | 100.0% | 100.0% | 100.0% | 75.0% | 65.0% | | Qwen2.5-3B | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | | DeepSeek-R1-Qwen3-8B-GGUF | 100.0% | 95.0% | 100.0% | 95.0% | 100.0% | | TinyLlama-1.1B | 100.0% | 90.0% | 100.0% | 70.0% | 100.0% | | Qwen2.5-1.5B | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | | Qwen2.5-0.5B | 60.0% | 65.0% | 67.5% | 95.0% | 65.0% | | SmolLM2-1.7B | 92.5% | 80.0% | 100.0% | 72.5% | 100.0% | | SmolLM2-360M | 100.0% | 100.0% | 100.0% | 45.0% | 100.0% | | Gemma-3-270M-GGUF | 85.0% | 97.5% | 90.0% | 87.5% | 100.0% | ## Per-domain exact-match breakdown (all models) | Model | CAL (n=2) | DB (n=27) | ECOMMERCE (n=8) | FILE (n=4) | HOME (n=7) | HTTP (n=25) | MQTT (n=15) | ROS (n=46) | VEHICLE (n=17) | WORKFLOW (n=49) | |---|---|---|---|---|---|---|---|---|---|---| | **[SAM]** SAM | 100.0% | 11.1% | 0.0% | 75.0% | 0.0% | 44.0% | 6.7% | 10.9% | 17.6% | 16.3% | | Qwen2.5-3B | 50.0% | 18.5% | 0.0% | 0.0% | 0.0% | 8.0% | 73.3% | 21.7% | 17.6% | 0.0% | | DeepSeek-R1-Qwen3-8B-GGUF | 50.0% | 0.0% | 0.0% | 25.0% | 0.0% | 8.0% | 73.3% | 10.9% | 17.6% | 0.0% | | TinyLlama-1.1B | 0.0% | 18.5% | 0.0% | 25.0% | 0.0% | 0.0% | 0.0% | 17.4% | 23.5% | 0.0% | | Qwen2.5-1.5B | 0.0% | 18.5% | 0.0% | 0.0% | 0.0% | 4.0% | 0.0% | 19.6% | 17.6% | 0.0% | | Qwen2.5-0.5B | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 4.3% | 17.6% | 0.0% | | SmolLM2-1.7B | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 4.3% | 17.6% | 0.0% | | SmolLM2-360M | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | | Gemma-3-270M-GGUF | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | ## How to read this report - The **headline aggregate** can be misleading because tier 1 is over-represented. - The **per-tier breakdown** is the honest comparison. Look at T1-T5 separately. - The **per-domain breakdown** shows which models are strong in which areas. - SAM is expected to win on the structured-output dimension (exact match on domains it was trained on) and to be **much smaller and faster**. Where general LLMs win is usually noisy/cross-domain (their broader training helps). - **No metric in isolation is the full story.** A reviewer should look at all three (exact match, latency, size) together to evaluate the trade-off. ## Reproducibility ```bash python sam_benchmark.py ``` Test suite regenerated deterministically from SEED=42. LLM weights from public HuggingFace repos. SAM weights from `AMFORGE/sam-v1-checkpoints/best.pt` (auto-downloaded if not present locally). Raw per-case results are in `results_full.json` (one entry per (model, case) including model output and pass/fail flags). Per-suite aggregates are in `results.json`. The test suite itself is in `test_suite.json`.