harrrshall's picture
Release BarunAction-35M candidate-v2
5a46e5d verified
|
Raw
History Blame Contribute Delete
4.45 kB
# Curated benchmark evidence
This directory contains the public, product-facing evidence for the BarunAction-35M candidate-v2
Mobile Actions development result. It intentionally excludes raw dataset rows, rendered prompts,
gold labels, official evaluation contents, infrastructure logs, and unrelated research history.
## Unpopulated sub-100M scaffold
[`sub100m/catalog.json`](sub100m/catalog.json) is an empty catalog template, not a completed model
survey or benchmark result. [`sub100m/lane-contracts.json`](sub100m/lane-contracts.json) freezes
eligibility plus separate off-the-shelf and matched-adaptation contracts before a researched roster
or predictions exist; it explicitly makes no exhaustive-coverage claim. Validate the scaffold with
`python scripts/validate_sub100m_benchmark.py`. Future predictions must be scored with the existing
`barunaction score-mobile` path and require a newly frozen, pre-outcome catalog version first.
## Evidence inventory
| File | Rows | Contents | SHA-256 |
| --- | ---: | --- | --- |
| [`barunaction-predictions.jsonl`](evidence/barunaction-predictions.jsonl) | 756 | IDs, raw candidate-v2 continuations, token counts, truncation/failure fields | `e5aea59e5090e3aa7c9dbd6b1a810c753410a159ce18dc8bb889fde3a58b8665` |
| [`barunaction-aggregate.json`](evidence/barunaction-aggregate.json) | 756 scored | Complete candidate-v2 aggregate | `5d7244d2fa449a7ce4b2aa3a20fc095fb118a38920b1dd9181f56669a7ddfea1` |
| [`qwen-predictions.jsonl`](evidence/qwen-predictions.jsonl) | 756 | IDs and raw Qwen continuations under the matched semantic recipe | `836788091d08b931bf37565d197c4053e6e6f4664d379c38e7668ae9b8739b08` |
| [`qwen-aggregate.json`](evidence/qwen-aggregate.json) | 756 scored | Complete Qwen aggregate | `2184b93d383f1136c050d95e71fb129bbe083b3fea4de86da6bdf5ec7f3e1756` |
| [`qwen-paired-outcomes.jsonl`](evidence/qwen-paired-outcomes.jsonl) | 756 | Per-ID candidate/baseline exactness and paired outcome | `ad5d44e741f9b98a0a0ffaadfa147c142cdaf8fea8801a11b62b7d323e93730a` |
| [`int8-paired-outcomes.jsonl`](evidence/int8-paired-outcomes.jsonl) | 756 | Per-ID float/int8 exactness and transition | `e63e48dc8149b615ab03a14fb9fff9b1435ab59c6623e72c16b44d34d7c651ad` |
| [`qwen-provenance.json`](evidence/qwen-provenance.json) | — | Exact model revision, files, architecture, and parameter audit | `1a1c19cdfacaa25e296ab1c3f92a42ebc8c08802c0cec8047ab581d50669769e` |
| [`manifest.json`](evidence/manifest.json) | — | File inventory and redistribution declarations | `d7d4223a58ac0277653c44209d9cb35edc0d98754b0c543019c0be67545bd318` |
Predictions are included so users can audit raw-output validity and compare model behavior. The
paired files expose sample-level correctness without copying prompts or gold actions. Aggregate
files are direct outputs of the checked-in deterministic scorer.
## Dataset boundary
The evidence uses `google/mobile-actions` revision
`e920309bc2acbc2e99a5e3201cf37df2b9fd9151`, licensed CC BY 4.0. The reported population is a
756-row grouped development split derived only from its public training rows. Its manifest hash is
`988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55`.
The manifest itself is not redistributed. Regenerate it locally with the pinned adapter following
[`docs/training.md`](../docs/training.md). The 961 official evaluation rows were never parsed,
materialized, or included here.
## Recompute
After regenerating and hash-checking the development manifest:
```console
python - <<'PY'
from barunlm.evaluation.mobile_actions import write_scores
write_scores(
"./local-data/mobile-actions/dev.jsonl",
"./benchmarks/evidence/barunaction-predictions.jsonl",
"./reproduced/barunaction",
)
write_scores(
"./local-data/mobile-actions/dev.jsonl",
"./benchmarks/evidence/qwen-predictions.jsonl",
"./reproduced/qwen",
)
PY
```
The exact parser, validator, and metric rules are in [`docs/evaluation.md`](../docs/evaluation.md).
The result table and comparison caveats are in [`docs/benchmark.md`](../docs/benchmark.md).
## Interpretation
Candidate-v2 scored 602/756 strict AST exact; the matched Qwen2.5-0.5B-Instruct baseline scored
663/756. BarunAction is 14.09 times smaller and retains 90.80% of Qwen's exact-match rate, but it
does not win this comparison. Both results are one-seed reused development measurements over an
all-`CALL` seven-tool population, not official-test or safety results.