Evaluation and scorer contract
BarunAction uses deterministic, model-free scoring. Raw model text is parsed as Action IR v1, validated against the supplied tool schemas, and compared with a canonical gold AST. No evaluator model, fuzzy string rule, output extraction, or repair step participates in the headline score.
The released versions are:
| Layer | Version |
|---|---|
| Action IR evaluator | action-ir-v1.0.0 |
| Mobile Actions scorer | barun-mobile-actions-score-v1 |
| Float generation record | barun-greedy-generation-v1 |
| Int8 generation record | barun-greedy-generation-int8-v1 |
Source of truth
| Concern | Implementation |
|---|---|
| Strict JSON decoding and Action IR validation | src/barunlm/evaluation/action_ir.py |
| Sample and aggregate metric logic | src/barunlm/evaluation/evaluator.py |
| Mobile tool-schema parsing and aligned scoring | src/barunlm/evaluation/mobile_actions.py |
| Deterministic generation and checkpoint verification | src/barunlm/evaluation/generation.py |
| Generate-and-score command | scripts/evaluate_mobile_actions.py |
| Scorer regression tests | tests/test_mobile_actions_scoring.py |
Documentation is explanatory; these checked-in functions define the executable metric.
Strict JSON parsing
decode_json_object() accepts exactly one UTF-8 JSON object. Leading and trailing JSON whitespace
is allowed. The following are deterministic parse failures:
- prose, Markdown fences, prefixes, suffixes, or a second JSON value;
- a non-object top-level value;
- duplicate object keys;
NaN,Infinity, numeric overflow, or non-finite values;- invalid UTF-8; and
- collection nesting deeper than 64 levels.
The evaluator never searches for a JSON substring and never retries with a repaired output.
Schema validation
validate_action_ir() requires one of four exact decision shapes:
{"calls":[{"args":{"query":"Cubbon Park"},"tool":"show_map"}],"decision":"CALL","mode":"SINGLE"}
{"calls":[{"args":{"body":"Running late","to":"Asha"},"tool":"send_message"}],"decision":"CONFIRM","mode":"SINGLE"}
{"decision":"CLARIFY","missing":["datetime"]}
{"decision":"ABSTAIN"}
Unknown or missing fields fail. CALL and CONFIRM require a non-empty calls list and one of
SINGLE, SERIAL, or PARALLEL. A call must contain exactly tool and args; the tool must be
declared; required arguments must be present; extra arguments are allowed only when the schema says
so; and every nested value must have the declared JSON type. SINGLE requires exactly one call.
CLARIFY.missing must be non-empty, sorted, unique, and string-valued. ABSTAIN may contain no
other field.
Canonical equality
After validation:
- object keys are sorted recursively;
- structural JSON whitespace and input key order are ignored;
- strings and object keys are Unicode NFC-normalized;
- whitespace inside string values remains exact;
- integers, numbers, booleans, nulls, arrays, and objects retain their JSON types;
- no optional argument default is inserted;
SINGLEandSERIALcalls remain ordered; andPARALLELcalls are compared as a multiset of canonical calls, preserving duplicates.
Two outputs are AST exact only when their decisions, modes, calls, and typed argument values match under those rules. Semantically similar dates, addresses, names, or strings do not receive credit unless they are represented identically after the declared deterministic canonicalization.
Mobile Actions adapter
For each manifest row, schemas_from_prompt() parses the exact TOOLS block rendered by
barun-mobile-actions-adapter-v2. It rejects malformed lines, duplicate tools, duplicate argument
names, and any type outside the frozen string, integer, number, or boolean subset. Tool
descriptions are model-visible but do not alter equality.
score_rows() requires a one-to-one ID join:
- manifest IDs and prediction IDs must each be unique;
- every manifest ID must have one prediction;
- no unknown prediction ID is allowed; and
- predictions may contain only
id,prediction_raw,truncated,generation_failure,prompt_tokens, andgenerated_tokens.
For this diagnostic, the deterministic execution hook reports success exactly when the Action IR
AST matches. Therefore policy-safe executable success and AST exact match have the same 602/756
value on the released all-CALL population. This equivalence is specific to this adapter and must
not be generalized to a real tool simulator.
Failure accounting
A row is incorrect when any of the following occurs:
- prediction is missing;
- generation reports a failure;
- generation reaches the 192-token cap and is marked truncated;
- raw text fails strict JSON parsing;
- parsed JSON fails Action IR or tool-schema validation; or
- the valid AST differs from gold.
Failures remain visible in failure_counts; they are not dropped from denominators. Candidate-v2
produced 756/756 parse-valid outputs, 755/756 schema-valid outputs, and 602/756 exact outputs. Its
single schema-invalid output used an invalid call count. Qwen produced 755/756 parse-valid,
754/756 schema-valid, and 663/756 exact outputs: one invalid JSON string and one additional parsed
object with an unknown field.
Secondary diagnostics
The aggregate also reports decision accuracy, per-tool metrics, argument-key and argument-value F1, per-scenario success, and failure categories.
Argument facts include sample ID, call index, tool, argument name, and—when scoring values—the
canonical JSON value. Including the sample ID prevents a correct fact on one row from canceling a
miss on another. PARALLEL calls are canonically sorted before assigning call indices. A row with
no gold and no predicted argument facts receives row-level macro-F1 1; errors on such rows remain
visible in decision and primary metrics.
The aggregate schema contains false-action, clarification, confirmation, and abstention fields, but the released 756-row population has zero eligible examples for those decisions. Values with a zero denominator are placeholders, not evidence of safe behavior.
Generation contract
The headline float and Qwen results use unconstrained deterministic greedy decoding:
do_sample=false;num_beams=1;temperature,top_k, andtop_pdisabled;- repetition penalty
1.0; - no-repeat n-gram size
0; and - maximum 192 new tokens.
There is no grammar constraint. Candidate-v2 used generation batch size 128 in its selection run; the matched Qwen lane used 64. Batch size is recorded because it can affect runtime, but it does not change the scoring definition.
Re-score the checked-in predictions
The repository distributes raw predictions and aggregate evidence, but not the prompt/gold
manifest. Regenerate the pinned development manifest following
docs/training.md, verify SHA-256
988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55, then run:
python - <<'PY'
from barunlm.evaluation.mobile_actions import write_scores
write_scores(
"./local-data/mobile-actions/dev.jsonl",
"./benchmarks/evidence/barunaction-predictions.jsonl",
"./reproduced/barunaction",
)
write_scores(
"./local-data/mobile-actions/dev.jsonl",
"./benchmarks/evidence/qwen-predictions.jsonl",
"./reproduced/qwen",
)
PY
Expected aggregate hashes:
| Output | SHA-256 |
|---|---|
| BarunAction aggregate | 5d7244d2fa449a7ce4b2aa3a20fc095fb118a38920b1dd9181f56669a7ddfea1 |
| Qwen aggregate | 2184b93d383f1136c050d95e71fb129bbe083b3fea4de86da6bdf5ec7f3e1756 |
To generate and score a new float checkpoint in one pass:
python scripts/evaluate_mobile_actions.py \
--checkpoint ./models/BarunAction-35M \
--manifest ./local-data/mobile-actions/dev.jsonl \
--manifest-sha256 988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55 \
--output ./new-evaluation \
--device cpu \
--batch-size 16 \
--max-new-tokens 192
The output directory must not already exist. The generator verifies the checkpoint before loading it and writes sample-level predictions before the scorer writes sample and aggregate records.
Interpretation boundary
The Mobile development result tests strict serialization, familiar-tool selection, and argument binding under one seven-tool schema family. It does not measure official test performance, abstention, ambiguity resolution, confirmation, false-action rate, prompt-injection robustness, renamed or unseen schemas, broad personal-action utility, or real-world execution safety.