File size: 8,791 Bytes
5a46e5d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
# BarunAction-35M benchmark report

The released BarunAction-35M candidate-v2 scored **602/756 (79.63%) strict Action IR AST exact**
on a grouped Mobile Actions development split. A matched Qwen2.5-0.5B-Instruct baseline scored
**663/756 (87.70%)**. BarunAction is 14.09 times smaller and retains 90.80% of Qwen's exact-match
rate, but it trails by 61 rows, or 8.07 percentage points.

This is a useful compact-model result, not a larger-model win. It is one-seed, reused public
development evidence over a narrow all-`CALL` population.

## Headline result

| Model | Exact parameters | Parse valid | Schema valid | Strict AST exact | Truncated |
| --- | ---: | ---: | ---: | ---: | ---: |
| **BarunAction-35M candidate-v2** | **35,072,768** | 756/756 | 755/756 | **602/756 (79.63%)** | 0/756 |
| Qwen2.5-0.5B-Instruct matched baseline | 494,032,768 | 755/756 | 754/756 | 663/756 (87.70%) | 0/756 |

The paired outcomes are:

| Outcome | Rows |
| --- | ---: |
| Both correct | 583 |
| BarunAction only | 19 |
| Qwen only | 80 |
| Both wrong | 74 |
| Total ties | 657 |

These counts can be inspected directly in
[`benchmarks/evidence/qwen-paired-outcomes.jsonl`](../benchmarks/evidence/qwen-paired-outcomes.jsonl).

## Population

- Dataset: `google/mobile-actions`.
- Revision: `e920309bc2acbc2e99a5e3201cf37df2b9fd9151`.
- License: CC BY 4.0.
- Training members: 7,937 rows derived from the source training split.
- Development members: 756 rows derived from the same source training split.
- Official evaluation members: 961 rows, never parsed or evaluated.
- Tools: calendar creation, contact creation, Wi-Fi settings, email, maps, and flashlight on/off.
- Gold decisions: all 756 are `CALL`; no row measures abstention, clarification, confirmation, or
  unsafe-request handling.

The development manifest SHA-256 is
`988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55`. It is not distributed;
users regenerate it from the pinned public source. See the [data card](barunaction-data-card.md).

## Metric

The primary metric is `action-ir-v1.0.0` AST exact match under
`barun-mobile-actions-score-v1`. A correct prediction must:

1. be exactly one strict JSON object;
2. validate against the row's rendered tool schemas;
3. match the gold `decision` and call `mode`;
4. match every tool name and typed argument value; and
5. preserve ordered calls for `SINGLE` and `SERIAL`, while matching `PARALLEL` calls as an
   order-independent multiset.

Object-key order and JSON structural whitespace are ignored after parsing. Strings are NFC
normalized but internal whitespace is preserved. There is no JSON extraction, Markdown removal,
type coercion, argument defaulting, or repair. Missing predictions, generation failures,
truncations, parse failures, and schema failures are incorrect.

The exact implementation and source paths are documented in [Evaluation](evaluation.md).

## Matched comparison design

The Qwen comparison used
`Qwen/Qwen2.5-0.5B-Instruct` revision
`7ae557604adf67be50417f59c2c2f167def9a775`, licensed Apache-2.0. Its checkpoint contains exactly
494,032,768 unique BF16 parameters across 290 tensors. It is about **0.494B parameters**, not 500B
and not 500 billion.

The two lanes matched:

- the exact 7,937 training IDs and 756 development IDs;
- semantic system content, ordered tool schemas, user text, and canonical targets;
- one full-parameter, response-only training pass;
- effective batch size 63, 126 optimizer steps, and seed 17;
- final-checkpoint-only scoring; and
- unconstrained deterministic greedy decoding with a 192-token generation cap.

They did not match upstream pretraining, prior instruction tuning, tokenizer-native control tokens,
tokenization, architecture, learning rate, memory technique, or total development-selection
budget. BarunAction used its native reserved-role prompt and `1e-4` peak learning rate. Qwen used
its native chat template, `2e-5`, per-device batch 21, accumulation 3, and gradient checkpointing.
BarunAction had already been selected from a three-trial development budget; Qwen used one frozen
recipe. The comparison is therefore matched semantic adaptation, not an assertion that parameter
count alone caused the difference.

Model identity and file hashes are in
[`benchmarks/evidence/qwen-provenance.json`](../benchmarks/evidence/qwen-provenance.json).
The provider-neutral matched-comparison record is
[`configs/benchmarks/qwen2.5-0.5b-matched.json`](../configs/benchmarks/qwen2.5-0.5b-matched.json),
SHA-256 `c40d8b0a370fd4e5ddfe20613ef31f00e81d2a5f941978cf3671009796dc608a`.

## Where the models differ

Scenario success uses the first gold call to name a row, so multi-call rows appear under their
first tool. It is diagnostic rather than a balanced benchmark.

| First-call scenario | Rows | BarunAction | Qwen |
| --- | ---: | ---: | ---: |
| `create_calendar_event` | 178 | 114 (64.04%) | 143 (80.34%) |
| `create_contact` | 128 | 104 (81.25%) | 118 (92.19%) |
| `open_wifi_settings` | 117 | 111 (94.87%) | 114 (97.44%) |
| `send_email` | 85 | 77 (90.59%) | 78 (91.76%) |
| `show_map` | 92 | 58 (63.04%) | 60 (65.22%) |
| `turn_off_flashlight` | 74 | 67 (90.54%) | 72 (97.30%) |
| `turn_on_flashlight` | 82 | 71 (86.59%) | 78 (95.12%) |

BarunAction selected tools accurately but lost more exact rows on argument values:

| Diagnostic | BarunAction | Qwen |
| --- | ---: | ---: |
| Argument-key micro-F1 | 99.73% | 99.65% |
| Argument-value micro-F1 | 90.87% | 94.52% |
| Argument-value macro-F1 | 90.33% | 92.74% |

The largest observed weakness is exact calendar and map argument binding. These error slices are
descriptive on the reused development population and must not be treated as a new tuning set.

## Candidate and int8 selection context

The earlier BarunAction candidate-v1 scored 578/756. The candidate-v2 selection compared two new
final checkpoints: `batch63` scored 602/756 and the competing `hardmix70` arm scored 566/756. The
602/756 result therefore selected candidate-v2 and is not an independent estimate.

The retained Darwin ARM64 dynamic-int8 derivative later scored 607/756 on the same IDs. Against
the frozen 602-row float reference it fixed eight rows and regressed three; against a same-host
FP32 control at 603/756 it fixed seven and regressed three. That post-selection retention check is
not another candidate-selection result and is not evidence that int8 generally improves accuracy.
See [Int8 quantization](int8-quantization.md).

## Curated evidence

The repository includes predictions, aggregates, paired correctness outcomes, and model provenance
without redistributing prompts or gold labels:

- [`barunaction-predictions.jsonl`](../benchmarks/evidence/barunaction-predictions.jsonl)
- [`barunaction-aggregate.json`](../benchmarks/evidence/barunaction-aggregate.json)
- [`qwen-predictions.jsonl`](../benchmarks/evidence/qwen-predictions.jsonl)
- [`qwen-aggregate.json`](../benchmarks/evidence/qwen-aggregate.json)
- [`qwen-paired-outcomes.jsonl`](../benchmarks/evidence/qwen-paired-outcomes.jsonl)
- [`int8-paired-outcomes.jsonl`](../benchmarks/evidence/int8-paired-outcomes.jsonl)
- [`manifest.json`](../benchmarks/evidence/manifest.json)

Each predictions file contains 756 IDs and raw model continuations. The paired files contain only
sample IDs and correctness transitions. Their hashes and redistribution declarations are bound in
the manifest. See [`benchmarks/README.md`](../benchmarks/README.md) for a field-level inventory.

## Reproduce the scores

First regenerate the development manifest as described in
[`docs/training.md`](training.md). Confirm its SHA-256 before scoring. Then recompute both aggregates
from the same local manifest:

```console
python - <<'PY'
from barunlm.evaluation.mobile_actions import write_scores

write_scores(
    "./local-data/mobile-actions/dev.jsonl",
    "./benchmarks/evidence/barunaction-predictions.jsonl",
    "./reproduced/barunaction",
)
write_scores(
    "./local-data/mobile-actions/dev.jsonl",
    "./benchmarks/evidence/qwen-predictions.jsonl",
    "./reproduced/qwen",
)
PY
```

The recomputed aggregates must match the checked-in files byte-for-byte. The scorer refuses
duplicate IDs, missing IDs, unknown IDs, unsupported prediction fields, malformed schemas, and an
existing output directory.

## Claim boundary

This benchmark does not establish safety, broad function calling, official Mobile Actions test
performance, unseen-schema behavior, statistical superiority, or production readiness. It has one
seed, no hidden-test uncertainty interval, a development-selected candidate, and no no-action
denominator. The truthful result is: **BarunAction delivers a substantial fraction of the matched
Qwen exact-match rate at 14.09 times fewer parameters, while Qwen remains 8.07 points ahead.**