File size: 8,895 Bytes
5a46e5d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
# Evaluation and scorer contract

BarunAction uses deterministic, model-free scoring. Raw model text is parsed as Action IR v1,
validated against the supplied tool schemas, and compared with a canonical gold AST. No evaluator
model, fuzzy string rule, output extraction, or repair step participates in the headline score.

The released versions are:

| Layer | Version |
| --- | --- |
| Action IR evaluator | `action-ir-v1.0.0` |
| Mobile Actions scorer | `barun-mobile-actions-score-v1` |
| Float generation record | `barun-greedy-generation-v1` |
| Int8 generation record | `barun-greedy-generation-int8-v1` |

## Source of truth

| Concern | Implementation |
| --- | --- |
| Strict JSON decoding and Action IR validation | [`src/barunlm/evaluation/action_ir.py`](../src/barunlm/evaluation/action_ir.py) |
| Sample and aggregate metric logic | [`src/barunlm/evaluation/evaluator.py`](../src/barunlm/evaluation/evaluator.py) |
| Mobile tool-schema parsing and aligned scoring | [`src/barunlm/evaluation/mobile_actions.py`](../src/barunlm/evaluation/mobile_actions.py) |
| Deterministic generation and checkpoint verification | [`src/barunlm/evaluation/generation.py`](../src/barunlm/evaluation/generation.py) |
| Generate-and-score command | [`scripts/evaluate_mobile_actions.py`](../scripts/evaluate_mobile_actions.py) |
| Scorer regression tests | [`tests/test_mobile_actions_scoring.py`](../tests/test_mobile_actions_scoring.py) |

Documentation is explanatory; these checked-in functions define the executable metric.

## Strict JSON parsing

`decode_json_object()` accepts exactly one UTF-8 JSON object. Leading and trailing JSON whitespace
is allowed. The following are deterministic parse failures:

- prose, Markdown fences, prefixes, suffixes, or a second JSON value;
- a non-object top-level value;
- duplicate object keys;
- `NaN`, `Infinity`, numeric overflow, or non-finite values;
- invalid UTF-8; and
- collection nesting deeper than 64 levels.

The evaluator never searches for a JSON substring and never retries with a repaired output.

## Schema validation

`validate_action_ir()` requires one of four exact decision shapes:

```json
{"calls":[{"args":{"query":"Cubbon Park"},"tool":"show_map"}],"decision":"CALL","mode":"SINGLE"}
```

```json
{"calls":[{"args":{"body":"Running late","to":"Asha"},"tool":"send_message"}],"decision":"CONFIRM","mode":"SINGLE"}
```

```json
{"decision":"CLARIFY","missing":["datetime"]}
```

```json
{"decision":"ABSTAIN"}
```

Unknown or missing fields fail. `CALL` and `CONFIRM` require a non-empty `calls` list and one of
`SINGLE`, `SERIAL`, or `PARALLEL`. A call must contain exactly `tool` and `args`; the tool must be
declared; required arguments must be present; extra arguments are allowed only when the schema says
so; and every nested value must have the declared JSON type. `SINGLE` requires exactly one call.

`CLARIFY.missing` must be non-empty, sorted, unique, and string-valued. `ABSTAIN` may contain no
other field.

## Canonical equality

After validation:

- object keys are sorted recursively;
- structural JSON whitespace and input key order are ignored;
- strings and object keys are Unicode NFC-normalized;
- whitespace inside string values remains exact;
- integers, numbers, booleans, nulls, arrays, and objects retain their JSON types;
- no optional argument default is inserted;
- `SINGLE` and `SERIAL` calls remain ordered; and
- `PARALLEL` calls are compared as a multiset of canonical calls, preserving duplicates.

Two outputs are AST exact only when their decisions, modes, calls, and typed argument values match
under those rules. Semantically similar dates, addresses, names, or strings do not receive credit
unless they are represented identically after the declared deterministic canonicalization.

## Mobile Actions adapter

For each manifest row, `schemas_from_prompt()` parses the exact `TOOLS` block rendered by
`barun-mobile-actions-adapter-v2`. It rejects malformed lines, duplicate tools, duplicate argument
names, and any type outside the frozen `string`, `integer`, `number`, or `boolean` subset. Tool
descriptions are model-visible but do not alter equality.

`score_rows()` requires a one-to-one ID join:

- manifest IDs and prediction IDs must each be unique;
- every manifest ID must have one prediction;
- no unknown prediction ID is allowed; and
- predictions may contain only `id`, `prediction_raw`, `truncated`, `generation_failure`,
  `prompt_tokens`, and `generated_tokens`.

For this diagnostic, the deterministic execution hook reports success exactly when the Action IR
AST matches. Therefore policy-safe executable success and AST exact match have the same 602/756
value on the released all-`CALL` population. This equivalence is specific to this adapter and must
not be generalized to a real tool simulator.

## Failure accounting

A row is incorrect when any of the following occurs:

- prediction is missing;
- generation reports a failure;
- generation reaches the 192-token cap and is marked truncated;
- raw text fails strict JSON parsing;
- parsed JSON fails Action IR or tool-schema validation; or
- the valid AST differs from gold.

Failures remain visible in `failure_counts`; they are not dropped from denominators. Candidate-v2
produced 756/756 parse-valid outputs, 755/756 schema-valid outputs, and 602/756 exact outputs. Its
single schema-invalid output used an invalid call count. Qwen produced 755/756 parse-valid,
754/756 schema-valid, and 663/756 exact outputs: one invalid JSON string and one additional parsed
object with an unknown field.

## Secondary diagnostics

The aggregate also reports decision accuracy, per-tool metrics, argument-key and argument-value
F1, per-scenario success, and failure categories.

Argument facts include sample ID, call index, tool, argument name, and—when scoring values—the
canonical JSON value. Including the sample ID prevents a correct fact on one row from canceling a
miss on another. `PARALLEL` calls are canonically sorted before assigning call indices. A row with
no gold and no predicted argument facts receives row-level macro-F1 1; errors on such rows remain
visible in decision and primary metrics.

The aggregate schema contains false-action, clarification, confirmation, and abstention fields,
but the released 756-row population has zero eligible examples for those decisions. Values with a
zero denominator are placeholders, not evidence of safe behavior.

## Generation contract

The headline float and Qwen results use unconstrained deterministic greedy decoding:

- `do_sample=false`;
- `num_beams=1`;
- `temperature`, `top_k`, and `top_p` disabled;
- repetition penalty `1.0`;
- no-repeat n-gram size `0`; and
- maximum 192 new tokens.

There is no grammar constraint. Candidate-v2 used generation batch size 128 in its selection run;
the matched Qwen lane used 64. Batch size is recorded because it can affect runtime, but it does
not change the scoring definition.

## Re-score the checked-in predictions

The repository distributes raw predictions and aggregate evidence, but not the prompt/gold
manifest. Regenerate the pinned development manifest following
[`docs/training.md`](training.md), verify SHA-256
`988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55`, then run:

```console
python - <<'PY'
from barunlm.evaluation.mobile_actions import write_scores

write_scores(
    "./local-data/mobile-actions/dev.jsonl",
    "./benchmarks/evidence/barunaction-predictions.jsonl",
    "./reproduced/barunaction",
)
write_scores(
    "./local-data/mobile-actions/dev.jsonl",
    "./benchmarks/evidence/qwen-predictions.jsonl",
    "./reproduced/qwen",
)
PY
```

Expected aggregate hashes:

| Output | SHA-256 |
| --- | --- |
| BarunAction aggregate | `5d7244d2fa449a7ce4b2aa3a20fc095fb118a38920b1dd9181f56669a7ddfea1` |
| Qwen aggregate | `2184b93d383f1136c050d95e71fb129bbe083b3fea4de86da6bdf5ec7f3e1756` |

To generate and score a new float checkpoint in one pass:

```console
python scripts/evaluate_mobile_actions.py \
  --checkpoint ./models/BarunAction-35M \
  --manifest ./local-data/mobile-actions/dev.jsonl \
  --manifest-sha256 988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55 \
  --output ./new-evaluation \
  --device cpu \
  --batch-size 16 \
  --max-new-tokens 192
```

The output directory must not already exist. The generator verifies the checkpoint before loading
it and writes sample-level predictions before the scorer writes sample and aggregate records.

## Interpretation boundary

The Mobile development result tests strict serialization, familiar-tool selection, and argument
binding under one seven-tool schema family. It does not measure official test performance,
abstention, ambiguity resolution, confirmation, false-action rate, prompt-injection robustness,
renamed or unseen schemas, broad personal-action utility, or real-world execution safety.