harrrshall's picture
Release BarunAction-35M candidate-v2
5a46e5d verified
|
Raw
History Blame Contribute Delete
4.45 kB

Curated benchmark evidence

This directory contains the public, product-facing evidence for the BarunAction-35M candidate-v2 Mobile Actions development result. It intentionally excludes raw dataset rows, rendered prompts, gold labels, official evaluation contents, infrastructure logs, and unrelated research history.

Unpopulated sub-100M scaffold

sub100m/catalog.json is an empty catalog template, not a completed model survey or benchmark result. sub100m/lane-contracts.json freezes eligibility plus separate off-the-shelf and matched-adaptation contracts before a researched roster or predictions exist; it explicitly makes no exhaustive-coverage claim. Validate the scaffold with python scripts/validate_sub100m_benchmark.py. Future predictions must be scored with the existing barunaction score-mobile path and require a newly frozen, pre-outcome catalog version first.

Evidence inventory

File Rows Contents SHA-256
barunaction-predictions.jsonl 756 IDs, raw candidate-v2 continuations, token counts, truncation/failure fields e5aea59e5090e3aa7c9dbd6b1a810c753410a159ce18dc8bb889fde3a58b8665
barunaction-aggregate.json 756 scored Complete candidate-v2 aggregate 5d7244d2fa449a7ce4b2aa3a20fc095fb118a38920b1dd9181f56669a7ddfea1
qwen-predictions.jsonl 756 IDs and raw Qwen continuations under the matched semantic recipe 836788091d08b931bf37565d197c4053e6e6f4664d379c38e7668ae9b8739b08
qwen-aggregate.json 756 scored Complete Qwen aggregate 2184b93d383f1136c050d95e71fb129bbe083b3fea4de86da6bdf5ec7f3e1756
qwen-paired-outcomes.jsonl 756 Per-ID candidate/baseline exactness and paired outcome ad5d44e741f9b98a0a0ffaadfa147c142cdaf8fea8801a11b62b7d323e93730a
int8-paired-outcomes.jsonl 756 Per-ID float/int8 exactness and transition e63e48dc8149b615ab03a14fb9fff9b1435ab59c6623e72c16b44d34d7c651ad
qwen-provenance.json Exact model revision, files, architecture, and parameter audit 1a1c19cdfacaa25e296ab1c3f92a42ebc8c08802c0cec8047ab581d50669769e
manifest.json File inventory and redistribution declarations d7d4223a58ac0277653c44209d9cb35edc0d98754b0c543019c0be67545bd318

Predictions are included so users can audit raw-output validity and compare model behavior. The paired files expose sample-level correctness without copying prompts or gold actions. Aggregate files are direct outputs of the checked-in deterministic scorer.

Dataset boundary

The evidence uses google/mobile-actions revision e920309bc2acbc2e99a5e3201cf37df2b9fd9151, licensed CC BY 4.0. The reported population is a 756-row grouped development split derived only from its public training rows. Its manifest hash is 988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55.

The manifest itself is not redistributed. Regenerate it locally with the pinned adapter following docs/training.md. The 961 official evaluation rows were never parsed, materialized, or included here.

Recompute

After regenerating and hash-checking the development manifest:

python - <<'PY'
from barunlm.evaluation.mobile_actions import write_scores

write_scores(
    "./local-data/mobile-actions/dev.jsonl",
    "./benchmarks/evidence/barunaction-predictions.jsonl",
    "./reproduced/barunaction",
)
write_scores(
    "./local-data/mobile-actions/dev.jsonl",
    "./benchmarks/evidence/qwen-predictions.jsonl",
    "./reproduced/qwen",
)
PY

The exact parser, validator, and metric rules are in docs/evaluation.md. The result table and comparison caveats are in docs/benchmark.md.

Interpretation

Candidate-v2 scored 602/756 strict AST exact; the matched Qwen2.5-0.5B-Instruct baseline scored 663/756. BarunAction is 14.09 times smaller and retains 90.80% of Qwen's exact-match rate, but it does not win this comparison. Both results are one-seed reused development measurements over an all-CALL seven-tool population, not official-test or safety results.