| # BarunAction-35M benchmark report |
|
|
| The released BarunAction-35M candidate-v2 scored **602/756 (79.63%) strict Action IR AST exact** |
| on a grouped Mobile Actions development split. A matched Qwen2.5-0.5B-Instruct baseline scored |
| **663/756 (87.70%)**. BarunAction is 14.09 times smaller and retains 90.80% of Qwen's exact-match |
| rate, but it trails by 61 rows, or 8.07 percentage points. |
|
|
| This is a useful compact-model result, not a larger-model win. It is one-seed, reused public |
| development evidence over a narrow all-`CALL` population. |
|
|
| ## Headline result |
|
|
| | Model | Exact parameters | Parse valid | Schema valid | Strict AST exact | Truncated | |
| | --- | ---: | ---: | ---: | ---: | ---: | |
| | **BarunAction-35M candidate-v2** | **35,072,768** | 756/756 | 755/756 | **602/756 (79.63%)** | 0/756 | |
| | Qwen2.5-0.5B-Instruct matched baseline | 494,032,768 | 755/756 | 754/756 | 663/756 (87.70%) | 0/756 | |
|
|
| The paired outcomes are: |
|
|
| | Outcome | Rows | |
| | --- | ---: | |
| | Both correct | 583 | |
| | BarunAction only | 19 | |
| | Qwen only | 80 | |
| | Both wrong | 74 | |
| | Total ties | 657 | |
|
|
| These counts can be inspected directly in |
| [`benchmarks/evidence/qwen-paired-outcomes.jsonl`](../benchmarks/evidence/qwen-paired-outcomes.jsonl). |
|
|
| ## Population |
|
|
| - Dataset: `google/mobile-actions`. |
| - Revision: `e920309bc2acbc2e99a5e3201cf37df2b9fd9151`. |
| - License: CC BY 4.0. |
| - Training members: 7,937 rows derived from the source training split. |
| - Development members: 756 rows derived from the same source training split. |
| - Official evaluation members: 961 rows, never parsed or evaluated. |
| - Tools: calendar creation, contact creation, Wi-Fi settings, email, maps, and flashlight on/off. |
| - Gold decisions: all 756 are `CALL`; no row measures abstention, clarification, confirmation, or |
| unsafe-request handling. |
|
|
| The development manifest SHA-256 is |
| `988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55`. It is not distributed; |
| users regenerate it from the pinned public source. See the [data card](barunaction-data-card.md). |
|
|
| ## Metric |
|
|
| The primary metric is `action-ir-v1.0.0` AST exact match under |
| `barun-mobile-actions-score-v1`. A correct prediction must: |
|
|
| 1. be exactly one strict JSON object; |
| 2. validate against the row's rendered tool schemas; |
| 3. match the gold `decision` and call `mode`; |
| 4. match every tool name and typed argument value; and |
| 5. preserve ordered calls for `SINGLE` and `SERIAL`, while matching `PARALLEL` calls as an |
| order-independent multiset. |
|
|
| Object-key order and JSON structural whitespace are ignored after parsing. Strings are NFC |
| normalized but internal whitespace is preserved. There is no JSON extraction, Markdown removal, |
| type coercion, argument defaulting, or repair. Missing predictions, generation failures, |
| truncations, parse failures, and schema failures are incorrect. |
|
|
| The exact implementation and source paths are documented in [Evaluation](evaluation.md). |
|
|
| ## Matched comparison design |
|
|
| The Qwen comparison used |
| `Qwen/Qwen2.5-0.5B-Instruct` revision |
| `7ae557604adf67be50417f59c2c2f167def9a775`, licensed Apache-2.0. Its checkpoint contains exactly |
| 494,032,768 unique BF16 parameters across 290 tensors. It is about **0.494B parameters**, not 500B |
| and not 500 billion. |
|
|
| The two lanes matched: |
|
|
| - the exact 7,937 training IDs and 756 development IDs; |
| - semantic system content, ordered tool schemas, user text, and canonical targets; |
| - one full-parameter, response-only training pass; |
| - effective batch size 63, 126 optimizer steps, and seed 17; |
| - final-checkpoint-only scoring; and |
| - unconstrained deterministic greedy decoding with a 192-token generation cap. |
|
|
| They did not match upstream pretraining, prior instruction tuning, tokenizer-native control tokens, |
| tokenization, architecture, learning rate, memory technique, or total development-selection |
| budget. BarunAction used its native reserved-role prompt and `1e-4` peak learning rate. Qwen used |
| its native chat template, `2e-5`, per-device batch 21, accumulation 3, and gradient checkpointing. |
| BarunAction had already been selected from a three-trial development budget; Qwen used one frozen |
| recipe. The comparison is therefore matched semantic adaptation, not an assertion that parameter |
| count alone caused the difference. |
|
|
| Model identity and file hashes are in |
| [`benchmarks/evidence/qwen-provenance.json`](../benchmarks/evidence/qwen-provenance.json). |
| The provider-neutral matched-comparison record is |
| [`configs/benchmarks/qwen2.5-0.5b-matched.json`](../configs/benchmarks/qwen2.5-0.5b-matched.json), |
| SHA-256 `c40d8b0a370fd4e5ddfe20613ef31f00e81d2a5f941978cf3671009796dc608a`. |
|
|
| ## Where the models differ |
|
|
| Scenario success uses the first gold call to name a row, so multi-call rows appear under their |
| first tool. It is diagnostic rather than a balanced benchmark. |
|
|
| | First-call scenario | Rows | BarunAction | Qwen | |
| | --- | ---: | ---: | ---: | |
| | `create_calendar_event` | 178 | 114 (64.04%) | 143 (80.34%) | |
| | `create_contact` | 128 | 104 (81.25%) | 118 (92.19%) | |
| | `open_wifi_settings` | 117 | 111 (94.87%) | 114 (97.44%) | |
| | `send_email` | 85 | 77 (90.59%) | 78 (91.76%) | |
| | `show_map` | 92 | 58 (63.04%) | 60 (65.22%) | |
| | `turn_off_flashlight` | 74 | 67 (90.54%) | 72 (97.30%) | |
| | `turn_on_flashlight` | 82 | 71 (86.59%) | 78 (95.12%) | |
|
|
| BarunAction selected tools accurately but lost more exact rows on argument values: |
|
|
| | Diagnostic | BarunAction | Qwen | |
| | --- | ---: | ---: | |
| | Argument-key micro-F1 | 99.73% | 99.65% | |
| | Argument-value micro-F1 | 90.87% | 94.52% | |
| | Argument-value macro-F1 | 90.33% | 92.74% | |
|
|
| The largest observed weakness is exact calendar and map argument binding. These error slices are |
| descriptive on the reused development population and must not be treated as a new tuning set. |
|
|
| ## Candidate and int8 selection context |
|
|
| The earlier BarunAction candidate-v1 scored 578/756. The candidate-v2 selection compared two new |
| final checkpoints: `batch63` scored 602/756 and the competing `hardmix70` arm scored 566/756. The |
| 602/756 result therefore selected candidate-v2 and is not an independent estimate. |
|
|
| The retained Darwin ARM64 dynamic-int8 derivative later scored 607/756 on the same IDs. Against |
| the frozen 602-row float reference it fixed eight rows and regressed three; against a same-host |
| FP32 control at 603/756 it fixed seven and regressed three. That post-selection retention check is |
| not another candidate-selection result and is not evidence that int8 generally improves accuracy. |
| See [Int8 quantization](int8-quantization.md). |
|
|
| ## Curated evidence |
|
|
| The repository includes predictions, aggregates, paired correctness outcomes, and model provenance |
| without redistributing prompts or gold labels: |
|
|
| - [`barunaction-predictions.jsonl`](../benchmarks/evidence/barunaction-predictions.jsonl) |
| - [`barunaction-aggregate.json`](../benchmarks/evidence/barunaction-aggregate.json) |
| - [`qwen-predictions.jsonl`](../benchmarks/evidence/qwen-predictions.jsonl) |
| - [`qwen-aggregate.json`](../benchmarks/evidence/qwen-aggregate.json) |
| - [`qwen-paired-outcomes.jsonl`](../benchmarks/evidence/qwen-paired-outcomes.jsonl) |
| - [`int8-paired-outcomes.jsonl`](../benchmarks/evidence/int8-paired-outcomes.jsonl) |
| - [`manifest.json`](../benchmarks/evidence/manifest.json) |
|
|
| Each predictions file contains 756 IDs and raw model continuations. The paired files contain only |
| sample IDs and correctness transitions. Their hashes and redistribution declarations are bound in |
| the manifest. See [`benchmarks/README.md`](../benchmarks/README.md) for a field-level inventory. |
|
|
| ## Reproduce the scores |
|
|
| First regenerate the development manifest as described in |
| [`docs/training.md`](training.md). Confirm its SHA-256 before scoring. Then recompute both aggregates |
| from the same local manifest: |
|
|
| ```console |
| python - <<'PY' |
| from barunlm.evaluation.mobile_actions import write_scores |
| |
| write_scores( |
| "./local-data/mobile-actions/dev.jsonl", |
| "./benchmarks/evidence/barunaction-predictions.jsonl", |
| "./reproduced/barunaction", |
| ) |
| write_scores( |
| "./local-data/mobile-actions/dev.jsonl", |
| "./benchmarks/evidence/qwen-predictions.jsonl", |
| "./reproduced/qwen", |
| ) |
| PY |
| ``` |
|
|
| The recomputed aggregates must match the checked-in files byte-for-byte. The scorer refuses |
| duplicate IDs, missing IDs, unknown IDs, unsupported prediction fields, malformed schemas, and an |
| existing output directory. |
|
|
| ## Claim boundary |
|
|
| This benchmark does not establish safety, broad function calling, official Mobile Actions test |
| performance, unseen-schema behavior, statistical superiority, or production readiness. It has one |
| seed, no hidden-test uncertainty interval, a development-selected candidate, and no no-action |
| denominator. The truthful result is: **BarunAction delivers a substantial fraction of the matched |
| Qwen exact-match rate at 14.09 times fewer parameters, while Qwen remains 8.07 points ahead.** |
|
|