BarunAction-35M / source /docs /evaluation.md
harrrshall's picture
Release BarunAction-35M candidate-v2
5a46e5d verified
|
Raw
History Blame Contribute Delete
8.9 kB

Evaluation and scorer contract

BarunAction uses deterministic, model-free scoring. Raw model text is parsed as Action IR v1, validated against the supplied tool schemas, and compared with a canonical gold AST. No evaluator model, fuzzy string rule, output extraction, or repair step participates in the headline score.

The released versions are:

Layer Version
Action IR evaluator action-ir-v1.0.0
Mobile Actions scorer barun-mobile-actions-score-v1
Float generation record barun-greedy-generation-v1
Int8 generation record barun-greedy-generation-int8-v1

Source of truth

Concern Implementation
Strict JSON decoding and Action IR validation src/barunlm/evaluation/action_ir.py
Sample and aggregate metric logic src/barunlm/evaluation/evaluator.py
Mobile tool-schema parsing and aligned scoring src/barunlm/evaluation/mobile_actions.py
Deterministic generation and checkpoint verification src/barunlm/evaluation/generation.py
Generate-and-score command scripts/evaluate_mobile_actions.py
Scorer regression tests tests/test_mobile_actions_scoring.py

Documentation is explanatory; these checked-in functions define the executable metric.

Strict JSON parsing

decode_json_object() accepts exactly one UTF-8 JSON object. Leading and trailing JSON whitespace is allowed. The following are deterministic parse failures:

  • prose, Markdown fences, prefixes, suffixes, or a second JSON value;
  • a non-object top-level value;
  • duplicate object keys;
  • NaN, Infinity, numeric overflow, or non-finite values;
  • invalid UTF-8; and
  • collection nesting deeper than 64 levels.

The evaluator never searches for a JSON substring and never retries with a repaired output.

Schema validation

validate_action_ir() requires one of four exact decision shapes:

{"calls":[{"args":{"query":"Cubbon Park"},"tool":"show_map"}],"decision":"CALL","mode":"SINGLE"}
{"calls":[{"args":{"body":"Running late","to":"Asha"},"tool":"send_message"}],"decision":"CONFIRM","mode":"SINGLE"}
{"decision":"CLARIFY","missing":["datetime"]}
{"decision":"ABSTAIN"}

Unknown or missing fields fail. CALL and CONFIRM require a non-empty calls list and one of SINGLE, SERIAL, or PARALLEL. A call must contain exactly tool and args; the tool must be declared; required arguments must be present; extra arguments are allowed only when the schema says so; and every nested value must have the declared JSON type. SINGLE requires exactly one call.

CLARIFY.missing must be non-empty, sorted, unique, and string-valued. ABSTAIN may contain no other field.

Canonical equality

After validation:

  • object keys are sorted recursively;
  • structural JSON whitespace and input key order are ignored;
  • strings and object keys are Unicode NFC-normalized;
  • whitespace inside string values remains exact;
  • integers, numbers, booleans, nulls, arrays, and objects retain their JSON types;
  • no optional argument default is inserted;
  • SINGLE and SERIAL calls remain ordered; and
  • PARALLEL calls are compared as a multiset of canonical calls, preserving duplicates.

Two outputs are AST exact only when their decisions, modes, calls, and typed argument values match under those rules. Semantically similar dates, addresses, names, or strings do not receive credit unless they are represented identically after the declared deterministic canonicalization.

Mobile Actions adapter

For each manifest row, schemas_from_prompt() parses the exact TOOLS block rendered by barun-mobile-actions-adapter-v2. It rejects malformed lines, duplicate tools, duplicate argument names, and any type outside the frozen string, integer, number, or boolean subset. Tool descriptions are model-visible but do not alter equality.

score_rows() requires a one-to-one ID join:

  • manifest IDs and prediction IDs must each be unique;
  • every manifest ID must have one prediction;
  • no unknown prediction ID is allowed; and
  • predictions may contain only id, prediction_raw, truncated, generation_failure, prompt_tokens, and generated_tokens.

For this diagnostic, the deterministic execution hook reports success exactly when the Action IR AST matches. Therefore policy-safe executable success and AST exact match have the same 602/756 value on the released all-CALL population. This equivalence is specific to this adapter and must not be generalized to a real tool simulator.

Failure accounting

A row is incorrect when any of the following occurs:

  • prediction is missing;
  • generation reports a failure;
  • generation reaches the 192-token cap and is marked truncated;
  • raw text fails strict JSON parsing;
  • parsed JSON fails Action IR or tool-schema validation; or
  • the valid AST differs from gold.

Failures remain visible in failure_counts; they are not dropped from denominators. Candidate-v2 produced 756/756 parse-valid outputs, 755/756 schema-valid outputs, and 602/756 exact outputs. Its single schema-invalid output used an invalid call count. Qwen produced 755/756 parse-valid, 754/756 schema-valid, and 663/756 exact outputs: one invalid JSON string and one additional parsed object with an unknown field.

Secondary diagnostics

The aggregate also reports decision accuracy, per-tool metrics, argument-key and argument-value F1, per-scenario success, and failure categories.

Argument facts include sample ID, call index, tool, argument name, and—when scoring values—the canonical JSON value. Including the sample ID prevents a correct fact on one row from canceling a miss on another. PARALLEL calls are canonically sorted before assigning call indices. A row with no gold and no predicted argument facts receives row-level macro-F1 1; errors on such rows remain visible in decision and primary metrics.

The aggregate schema contains false-action, clarification, confirmation, and abstention fields, but the released 756-row population has zero eligible examples for those decisions. Values with a zero denominator are placeholders, not evidence of safe behavior.

Generation contract

The headline float and Qwen results use unconstrained deterministic greedy decoding:

  • do_sample=false;
  • num_beams=1;
  • temperature, top_k, and top_p disabled;
  • repetition penalty 1.0;
  • no-repeat n-gram size 0; and
  • maximum 192 new tokens.

There is no grammar constraint. Candidate-v2 used generation batch size 128 in its selection run; the matched Qwen lane used 64. Batch size is recorded because it can affect runtime, but it does not change the scoring definition.

Re-score the checked-in predictions

The repository distributes raw predictions and aggregate evidence, but not the prompt/gold manifest. Regenerate the pinned development manifest following docs/training.md, verify SHA-256 988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55, then run:

python - <<'PY'
from barunlm.evaluation.mobile_actions import write_scores

write_scores(
    "./local-data/mobile-actions/dev.jsonl",
    "./benchmarks/evidence/barunaction-predictions.jsonl",
    "./reproduced/barunaction",
)
write_scores(
    "./local-data/mobile-actions/dev.jsonl",
    "./benchmarks/evidence/qwen-predictions.jsonl",
    "./reproduced/qwen",
)
PY

Expected aggregate hashes:

Output SHA-256
BarunAction aggregate 5d7244d2fa449a7ce4b2aa3a20fc095fb118a38920b1dd9181f56669a7ddfea1
Qwen aggregate 2184b93d383f1136c050d95e71fb129bbe083b3fea4de86da6bdf5ec7f3e1756

To generate and score a new float checkpoint in one pass:

python scripts/evaluate_mobile_actions.py \
  --checkpoint ./models/BarunAction-35M \
  --manifest ./local-data/mobile-actions/dev.jsonl \
  --manifest-sha256 988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55 \
  --output ./new-evaluation \
  --device cpu \
  --batch-size 16 \
  --max-new-tokens 192

The output directory must not already exist. The generator verifies the checkpoint before loading it and writes sample-level predictions before the scorer writes sample and aggregate records.

Interpretation boundary

The Mobile development result tests strict serialization, familiar-tool selection, and argument binding under one seven-tool schema family. It does not measure official test performance, abstention, ambiguity resolution, confirmation, false-action rate, prompt-injection robustness, renamed or unseen schemas, broad personal-action utility, or real-world execution safety.