BarunAction-35M / source /docs /evaluation.md
harrrshall's picture
Release BarunAction-35M candidate-v2
5a46e5d verified
|
Raw
History Blame Contribute Delete
8.9 kB
# Evaluation and scorer contract
BarunAction uses deterministic, model-free scoring. Raw model text is parsed as Action IR v1,
validated against the supplied tool schemas, and compared with a canonical gold AST. No evaluator
model, fuzzy string rule, output extraction, or repair step participates in the headline score.
The released versions are:
| Layer | Version |
| --- | --- |
| Action IR evaluator | `action-ir-v1.0.0` |
| Mobile Actions scorer | `barun-mobile-actions-score-v1` |
| Float generation record | `barun-greedy-generation-v1` |
| Int8 generation record | `barun-greedy-generation-int8-v1` |
## Source of truth
| Concern | Implementation |
| --- | --- |
| Strict JSON decoding and Action IR validation | [`src/barunlm/evaluation/action_ir.py`](../src/barunlm/evaluation/action_ir.py) |
| Sample and aggregate metric logic | [`src/barunlm/evaluation/evaluator.py`](../src/barunlm/evaluation/evaluator.py) |
| Mobile tool-schema parsing and aligned scoring | [`src/barunlm/evaluation/mobile_actions.py`](../src/barunlm/evaluation/mobile_actions.py) |
| Deterministic generation and checkpoint verification | [`src/barunlm/evaluation/generation.py`](../src/barunlm/evaluation/generation.py) |
| Generate-and-score command | [`scripts/evaluate_mobile_actions.py`](../scripts/evaluate_mobile_actions.py) |
| Scorer regression tests | [`tests/test_mobile_actions_scoring.py`](../tests/test_mobile_actions_scoring.py) |
Documentation is explanatory; these checked-in functions define the executable metric.
## Strict JSON parsing
`decode_json_object()` accepts exactly one UTF-8 JSON object. Leading and trailing JSON whitespace
is allowed. The following are deterministic parse failures:
- prose, Markdown fences, prefixes, suffixes, or a second JSON value;
- a non-object top-level value;
- duplicate object keys;
- `NaN`, `Infinity`, numeric overflow, or non-finite values;
- invalid UTF-8; and
- collection nesting deeper than 64 levels.
The evaluator never searches for a JSON substring and never retries with a repaired output.
## Schema validation
`validate_action_ir()` requires one of four exact decision shapes:
```json
{"calls":[{"args":{"query":"Cubbon Park"},"tool":"show_map"}],"decision":"CALL","mode":"SINGLE"}
```
```json
{"calls":[{"args":{"body":"Running late","to":"Asha"},"tool":"send_message"}],"decision":"CONFIRM","mode":"SINGLE"}
```
```json
{"decision":"CLARIFY","missing":["datetime"]}
```
```json
{"decision":"ABSTAIN"}
```
Unknown or missing fields fail. `CALL` and `CONFIRM` require a non-empty `calls` list and one of
`SINGLE`, `SERIAL`, or `PARALLEL`. A call must contain exactly `tool` and `args`; the tool must be
declared; required arguments must be present; extra arguments are allowed only when the schema says
so; and every nested value must have the declared JSON type. `SINGLE` requires exactly one call.
`CLARIFY.missing` must be non-empty, sorted, unique, and string-valued. `ABSTAIN` may contain no
other field.
## Canonical equality
After validation:
- object keys are sorted recursively;
- structural JSON whitespace and input key order are ignored;
- strings and object keys are Unicode NFC-normalized;
- whitespace inside string values remains exact;
- integers, numbers, booleans, nulls, arrays, and objects retain their JSON types;
- no optional argument default is inserted;
- `SINGLE` and `SERIAL` calls remain ordered; and
- `PARALLEL` calls are compared as a multiset of canonical calls, preserving duplicates.
Two outputs are AST exact only when their decisions, modes, calls, and typed argument values match
under those rules. Semantically similar dates, addresses, names, or strings do not receive credit
unless they are represented identically after the declared deterministic canonicalization.
## Mobile Actions adapter
For each manifest row, `schemas_from_prompt()` parses the exact `TOOLS` block rendered by
`barun-mobile-actions-adapter-v2`. It rejects malformed lines, duplicate tools, duplicate argument
names, and any type outside the frozen `string`, `integer`, `number`, or `boolean` subset. Tool
descriptions are model-visible but do not alter equality.
`score_rows()` requires a one-to-one ID join:
- manifest IDs and prediction IDs must each be unique;
- every manifest ID must have one prediction;
- no unknown prediction ID is allowed; and
- predictions may contain only `id`, `prediction_raw`, `truncated`, `generation_failure`,
`prompt_tokens`, and `generated_tokens`.
For this diagnostic, the deterministic execution hook reports success exactly when the Action IR
AST matches. Therefore policy-safe executable success and AST exact match have the same 602/756
value on the released all-`CALL` population. This equivalence is specific to this adapter and must
not be generalized to a real tool simulator.
## Failure accounting
A row is incorrect when any of the following occurs:
- prediction is missing;
- generation reports a failure;
- generation reaches the 192-token cap and is marked truncated;
- raw text fails strict JSON parsing;
- parsed JSON fails Action IR or tool-schema validation; or
- the valid AST differs from gold.
Failures remain visible in `failure_counts`; they are not dropped from denominators. Candidate-v2
produced 756/756 parse-valid outputs, 755/756 schema-valid outputs, and 602/756 exact outputs. Its
single schema-invalid output used an invalid call count. Qwen produced 755/756 parse-valid,
754/756 schema-valid, and 663/756 exact outputs: one invalid JSON string and one additional parsed
object with an unknown field.
## Secondary diagnostics
The aggregate also reports decision accuracy, per-tool metrics, argument-key and argument-value
F1, per-scenario success, and failure categories.
Argument facts include sample ID, call index, tool, argument name, and—when scoring values—the
canonical JSON value. Including the sample ID prevents a correct fact on one row from canceling a
miss on another. `PARALLEL` calls are canonically sorted before assigning call indices. A row with
no gold and no predicted argument facts receives row-level macro-F1 1; errors on such rows remain
visible in decision and primary metrics.
The aggregate schema contains false-action, clarification, confirmation, and abstention fields,
but the released 756-row population has zero eligible examples for those decisions. Values with a
zero denominator are placeholders, not evidence of safe behavior.
## Generation contract
The headline float and Qwen results use unconstrained deterministic greedy decoding:
- `do_sample=false`;
- `num_beams=1`;
- `temperature`, `top_k`, and `top_p` disabled;
- repetition penalty `1.0`;
- no-repeat n-gram size `0`; and
- maximum 192 new tokens.
There is no grammar constraint. Candidate-v2 used generation batch size 128 in its selection run;
the matched Qwen lane used 64. Batch size is recorded because it can affect runtime, but it does
not change the scoring definition.
## Re-score the checked-in predictions
The repository distributes raw predictions and aggregate evidence, but not the prompt/gold
manifest. Regenerate the pinned development manifest following
[`docs/training.md`](training.md), verify SHA-256
`988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55`, then run:
```console
python - <<'PY'
from barunlm.evaluation.mobile_actions import write_scores
write_scores(
"./local-data/mobile-actions/dev.jsonl",
"./benchmarks/evidence/barunaction-predictions.jsonl",
"./reproduced/barunaction",
)
write_scores(
"./local-data/mobile-actions/dev.jsonl",
"./benchmarks/evidence/qwen-predictions.jsonl",
"./reproduced/qwen",
)
PY
```
Expected aggregate hashes:
| Output | SHA-256 |
| --- | --- |
| BarunAction aggregate | `5d7244d2fa449a7ce4b2aa3a20fc095fb118a38920b1dd9181f56669a7ddfea1` |
| Qwen aggregate | `2184b93d383f1136c050d95e71fb129bbe083b3fea4de86da6bdf5ec7f3e1756` |
To generate and score a new float checkpoint in one pass:
```console
python scripts/evaluate_mobile_actions.py \
--checkpoint ./models/BarunAction-35M \
--manifest ./local-data/mobile-actions/dev.jsonl \
--manifest-sha256 988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55 \
--output ./new-evaluation \
--device cpu \
--batch-size 16 \
--max-new-tokens 192
```
The output directory must not already exist. The generator verifies the checkpoint before loading
it and writes sample-level predictions before the scorer writes sample and aggregate records.
## Interpretation boundary
The Mobile development result tests strict serialization, familiar-tool selection, and argument
binding under one seven-tool schema family. It does not measure official test performance,
abstention, ambiguity resolution, confirmation, false-action rate, prompt-injection robustness,
renamed or unseen schemas, broad personal-action utility, or real-world execution safety.