sam-v1 / BENCHMARK.md
ameforge's picture
Upload BENCHMARK.md with huggingface_hub
b866d36 verified
|
Raw
History Blame Contribute Delete
7.52 kB

SAM Benchmark v1 β€” Full Report

Methodology

This benchmark evaluates structured action generation: given a natural language task and a JSON schema, produce a syntactically valid JSON sequence of operations that matches a reference answer.

Test suite: 200 prompts, deterministically generated from SEED=42. The exact prompts are saved to test_suite.json for reproducibility and audit.

Tier distribution (current suite):

  • Tier 1 β€” Atomic β€” single op: 40 cases (20.0%)
  • Tier 2 β€” Compound β€” 2-3 ops same domain: 40 cases (20.0%)
  • Tier 3 β€” Noisy β€” typos / casing / fillers: 40 cases (20.0%)
  • Tier 4 β€” Long chain β€” 4+ ops: 40 cases (20.0%)
  • Tier 5 β€” Cross-domain workflow: 40 cases (20.0%)

Domain distribution (current suite):

  • WORKFLOW: 49 cases (24.5%)
  • ROS: 46 cases (23.0%)
  • DB: 27 cases (13.5%)
  • HTTP: 25 cases (12.5%)
  • VEHICLE: 17 cases (8.5%)
  • MQTT: 15 cases (7.5%)
  • ECOMMERCE: 8 cases (4.0%)
  • HOME: 7 cases (3.5%)
  • FILE: 4 cases (2.0%)
  • CAL: 2 cases (1.0%)

Metrics:

  • JSON Valid: output parses as valid JSON ∈ [0,100]%
  • Op Correct: output is a list of dicts AND every op name matches the expected set ∈ [0,100]%
  • Exact Match: full structural match with reference, with float tolerance 0.02 for numerics ∈ [0,100]%
  • Latency: mean wall-clock time per prompt in milliseconds (single-thread, no batching)
  • Size: model file size on disk (MB)

Important disclosures:

  1. SAM is trained on the same domain distribution as this benchmark (multi-domain JSON ops). General-purpose LLMs are evaluated zero-shot. This is a deliberate comparison of "specialist 35M model" vs "generalist N-B model" on a structured task, NOT a claim of general superiority.
  2. Tier distribution is currently imbalanced toward Tier 1 (atomic). Aggregate exact-match numbers are dominated by easy cases. Always read the per-tier breakdown below for fair comparison.
  3. Domain distribution is currently imbalanced toward ROS (robotics). This reflects SAM's lineage from Foros (ROS-only). Per-domain breakdowns reveal each model's actual strength.
  4. All LLMs are loaded from publicly available HuggingFace weights with no fine-tuning. Prompt template is identical across all LLMs.

Headline results (aggregate across all 200 cases)

Rank Model Params Size (MB) JSON Valid Op Correct Exact Match Latency (ms)
1 [SAM] SAM 33.4M 127.4 88.0% 63.5% 18.0% 715.4
2 Qwen2.5-3B 3085.9M 5886.0 100.0% 82.5% 16.0% 2061.1
3 DeepSeek-R1-Qwen3-8B-GGUF 8000.0M 4794.9 98.0% 79.0% 11.5% 31737.9
4 TinyLlama-1.1B 1100.0M 2098.2 92.0% 55.5% 9.0% 4199.3
5 Qwen2.5-1.5B 1543.7M 2944.4 100.0% 78.5% 9.0% 1852.4
6 Qwen2.5-0.5B 494.0M 942.3 70.5% 30.0% 2.5% 1937.2
7 SmolLM2-1.7B 1711.4M 3264.2 89.0% 38.5% 2.5% 1694.9
8 SmolLM2-360M 361.8M 690.1 89.0% 0.0% 0.0% 2042.9
9 Llama-3.2-1B ? 0.0 0.0% 0.0% 0.0% 0.0
10 Gemma-2-2B ? 0.0 0.0% 0.0% 0.0% 0.0
11 Llama-3.2-3B ? 0.0 0.0% 0.0% 0.0% 0.0
12 Phi-3.5-mini-3.8B ? 0.0 0.0% 0.0% 0.0% 0.0
13 Llama-3.2-3B-uns ? 0.0 0.0% 0.0% 0.0% 0.0
14 Gemma-3-270M-GGUF 270.0M 241.4 92.0% 2.0% 0.0% 6322.5
15 Ministral-3-8B-GGUF ? 0.0 0.0% 0.0% 0.0% 0.0

Per-tier exact-match breakdown (all models)

This is the breakdown that matters for fair comparison. Each model's exact-match score is shown for each difficulty tier.

Model T1 atomic T2 compound T3 noisy T4 long T5 cross
[SAM] SAM 37.5% (n=40) 27.5% (n=40) 25.0% (n=40) 0.0% (n=40) 0.0% (n=40)
Qwen2.5-3B 40.0% (n=40) 12.5% (n=40) 27.5% (n=40) 0.0% (n=40) 0.0% (n=40)
DeepSeek-R1-Qwen3-8B-GGUF 27.5% (n=40) 7.5% (n=40) 22.5% (n=40) 0.0% (n=40) 0.0% (n=40)
TinyLlama-1.1B 27.5% (n=40) 12.5% (n=40) 5.0% (n=40) 0.0% (n=40) 0.0% (n=40)
Qwen2.5-1.5B 30.0% (n=40) 12.5% (n=40) 2.5% (n=40) 0.0% (n=40) 0.0% (n=40)
Qwen2.5-0.5B 12.5% (n=40) 0.0% (n=40) 0.0% (n=40) 0.0% (n=40) 0.0% (n=40)
SmolLM2-1.7B 12.5% (n=40) 0.0% (n=40) 0.0% (n=40) 0.0% (n=40) 0.0% (n=40)
SmolLM2-360M 0.0% (n=40) 0.0% (n=40) 0.0% (n=40) 0.0% (n=40) 0.0% (n=40)
Gemma-3-270M-GGUF 0.0% (n=40) 0.0% (n=40) 0.0% (n=40) 0.0% (n=40) 0.0% (n=40)

Per-tier JSON-valid breakdown (all models)

Model T1 atomic T2 compound T3 noisy T4 long T5 cross
[SAM] SAM 100.0% 100.0% 100.0% 75.0% 65.0%
Qwen2.5-3B 100.0% 100.0% 100.0% 100.0% 100.0%
DeepSeek-R1-Qwen3-8B-GGUF 100.0% 95.0% 100.0% 95.0% 100.0%
TinyLlama-1.1B 100.0% 90.0% 100.0% 70.0% 100.0%
Qwen2.5-1.5B 100.0% 100.0% 100.0% 100.0% 100.0%
Qwen2.5-0.5B 60.0% 65.0% 67.5% 95.0% 65.0%
SmolLM2-1.7B 92.5% 80.0% 100.0% 72.5% 100.0%
SmolLM2-360M 100.0% 100.0% 100.0% 45.0% 100.0%
Gemma-3-270M-GGUF 85.0% 97.5% 90.0% 87.5% 100.0%

Per-domain exact-match breakdown (all models)

Model CAL (n=2) DB (n=27) ECOMMERCE (n=8) FILE (n=4) HOME (n=7) HTTP (n=25) MQTT (n=15) ROS (n=46) VEHICLE (n=17) WORKFLOW (n=49)
[SAM] SAM 100.0% 11.1% 0.0% 75.0% 0.0% 44.0% 6.7% 10.9% 17.6% 16.3%
Qwen2.5-3B 50.0% 18.5% 0.0% 0.0% 0.0% 8.0% 73.3% 21.7% 17.6% 0.0%
DeepSeek-R1-Qwen3-8B-GGUF 50.0% 0.0% 0.0% 25.0% 0.0% 8.0% 73.3% 10.9% 17.6% 0.0%
TinyLlama-1.1B 0.0% 18.5% 0.0% 25.0% 0.0% 0.0% 0.0% 17.4% 23.5% 0.0%
Qwen2.5-1.5B 0.0% 18.5% 0.0% 0.0% 0.0% 4.0% 0.0% 19.6% 17.6% 0.0%
Qwen2.5-0.5B 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 4.3% 17.6% 0.0%
SmolLM2-1.7B 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 4.3% 17.6% 0.0%
SmolLM2-360M 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0%
Gemma-3-270M-GGUF 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0%

How to read this report

  • The headline aggregate can be misleading because tier 1 is over-represented.
  • The per-tier breakdown is the honest comparison. Look at T1-T5 separately.
  • The per-domain breakdown shows which models are strong in which areas.
  • SAM is expected to win on the structured-output dimension (exact match on domains it was trained on) and to be much smaller and faster. Where general LLMs win is usually noisy/cross-domain (their broader training helps).
  • No metric in isolation is the full story. A reviewer should look at all three (exact match, latency, size) together to evaluate the trade-off.

Reproducibility

python sam_benchmark.py

Test suite regenerated deterministically from SEED=42. LLM weights from public HuggingFace repos. SAM weights from AMFORGE/sam-v1-checkpoints/best.pt (auto-downloaded if not present locally).

Raw per-case results are in results_full.json (one entry per (model, case) including model output and pass/fail flags). Per-suite aggregates are in results.json. The test suite itself is in test_suite.json.