File size: 8,242 Bytes
5a46e5d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 | # Training BarunAction-35M candidate-v2
BarunAction-35M `candidate-v2` is the final checkpoint of one full-parameter,
response-only supervised fine-tuning epoch over 7,937 Mobile Actions examples. It starts from the
exact public BarunLM-35M base checkpoint; it is not continued from another action-tuned checkpoint
and does not use adapters, distillation, preference optimization, DPO, or reinforcement learning.
The machine-readable public record is
[`training/candidate-v2-recipe.json`](../training/candidate-v2-recipe.json). The per-step loss and
learning-rate trace is [`training/metrics.jsonl`](../training/metrics.jsonl). This document explains
how those fields map to the implementation.
Their SHA-256 values are
`436b15489334476af810fa74db01063a72a70f7aa61161692341adc1ca15e523` and
`24fcf2e0544b89c14e2fe462ffd7f8f5359547bd4bd7cce2d7082ee1e0eeca15`, respectively.
## Immutable inputs
### Base checkpoint
| Field | Value |
| --- | --- |
| Repository | `harrrshall/BarunLM-35M` |
| Revision | `ef3e483a9fd7d906ecf2a7929babeffaf82d1d16` |
| Parameters | 35,072,768 unique parameters |
| `barun_config.json` | `9b3a1d71baa95a198744d250f9629231738d942570b8685c44307fd83dd33565` |
| `model.safetensors` | `f2a7c88b9f2c2e3584809081407ab136795d82e30e89b730e007781c45d01447` |
| `tokenizer.json` | `70ded9605fccd09c2340ca7e225361eab0ae8b4dbbb0d6e26343ab5183979db6` |
The trainer verifies all three hashes before constructing the optimizer. Every model parameter is
trainable. Tied input/output embeddings are counted once for parameter accounting and remain tied.
### Mobile Actions population
| Artifact | Rows | Tokens including prompt and target | Supervised target tokens | SHA-256 |
| --- | ---: | ---: | ---: | --- |
| Training manifest | 7,937 | 2,719,132 | 721,254 | `131473ccb5bfb51cac0439b42159e72ec4c598025e50364a52b122b056c2e84e` |
| Development manifest | 756 | 259,452 | 68,873 | `988bdce5874d1f1a775feeb5ba2b58cd2bdc128f57e73cb9a63d535fae7c1d55` |
| Split audit | — | — | — | `dc756f97c0a7ef706ec8ffefe2d57cf7e16d75a6ccd932f906d16b6a6ee2f83c` |
The source is `google/mobile-actions` revision
`e920309bc2acbc2e99a5e3201cf37df2b9fd9151`; its `dataset.jsonl` SHA-256 is
`91d251ee958cfd295af6c4504c236a3a1ad19517de240c3bc680bacfcbf7e7d9`. It contains 8,693
source-training rows and 961 official evaluation rows. Only the source-training portion was parsed
to build the 7,937/756 grouped split. The official 961 rows remained opaque and were never
materialized as prompts or labels.
See the [data card](barunaction-data-card.md) for the derivation, leakage checks, license, and
redistribution boundary.
## Prompt and target construction
The deterministic adapter is implemented in
[`src/barunlm/datasets/mobile_actions.py`](../src/barunlm/datasets/mobile_actions.py). For each row it
renders:
```text
<bos><system>
ACTION_IR_V1
NOW <source timestamp>
TOOLS
<seven ordered typed tool declarations>
<user>
<request>
<assistant>
```
The target is one canonical Action IR v1 JSON object followed by exactly one `<eos>` token. Ordered
source calls map to `SINGLE` or `SERIAL`; the source has no dependency annotation from which to
infer `PARALLEL`. Null argument fields are omitted rather than converted to strings or defaults.
Response-only loss is literal in the implementation:
- prompt token labels are set to `-100` and do not contribute to cross-entropy;
- target tokens and the appended EOS token are supervised;
- padding labels are also `-100`;
- examples are tokenized without implicit special tokens; and
- overlength policy is `error`, so no row is silently truncated or dropped.
The relevant code is
[`src/barunlm/training/data.py`](../src/barunlm/training/data.py), especially
`tokenize_examples()` and `collate_sft()`.
## Exact optimization recipe
| Setting | Candidate-v2 value |
| --- | ---: |
| Objective | full-parameter response-only cross-entropy |
| Seed | 17 |
| Epochs | 1 |
| Examples presented | 7,937, once each |
| Batch size | 63 |
| Gradient accumulation | 1 |
| Optimizer steps | 126 |
| Optimizer | AdamW, non-fused |
| Peak learning rate | `1e-4` |
| Warmup | 12 linear steps |
| Decay | cosine to `1e-5` |
| Betas | `(0.9, 0.95)` |
| Epsilon | `1e-8` |
| Weight decay | `0.1` on rank-2-or-higher parameters; `0` otherwise |
| Gradient clipping | global norm `1.0` |
| Maximum sequence length | 2,048 |
| Development batch size | 128 |
| Training precision | CUDA BF16 autocast with FP32 model/optimizer state |
| Deterministic algorithms | enabled |
| Checkpoint used for scoring | unconditional final step 126 |
The first 12 updates linearly warm from `1e-4 / 12` to `1e-4`; the remaining schedule decays by
cosine to exactly `1e-5`. Development response loss was measured before training and after the
epoch, not used for within-run early stopping. It moved from `1.9527305807448436` to
`0.011631452274616101`. The recorded reference run completed on one NVIDIA H200 with PyTorch
2.13.0 and CUDA 13.0 in 30.21 seconds; that timing is environment-specific, not a deployment
latency claim.
The optimizer, schedule, checkpointing, and deterministic batching are implemented in
[`src/barunlm/training/trainer.py`](../src/barunlm/training/trainer.py). The entry point is
[`scripts/train_sft.py`](../scripts/train_sft.py).
## Candidate selection
Candidate-v2 was selected on the same 756-row development population used in the reported score.
The selection compared two preregistered one-epoch arms against an earlier 578/756 reference:
| Arm | Training view | Batch size | Steps | AST exact |
| --- | --- | ---: | ---: | ---: |
| `batch63` | all 7,937 rows once | 63 | 126 | **602/756** |
| `hardmix70` | fixed-size hard-example mix | 127 | 63 | 566/756 |
The rule required a unique winner at least one percentage point above the 578/756 reference and
used only each arm's final checkpoint. `batch63` passed and became candidate-v2. Because the same
development population selected the candidate, 602/756 is reused development evidence rather
than an independent test result.
## Reconstruct the local manifests
Raw and processed dataset rows are deliberately not committed. Download the pinned source into a
local cache and regenerate the manifests outside the repository:
```console
python -m barunlm.datasets.mobile_actions \
--download-pinned \
--output-dir ./local-data/mobile-actions \
--tokenizer-json ./models/BarunLM-35M/tokenizer.json \
--tokenizer-id harrrshall/BarunLM-35M \
--tokenizer-revision ef3e483a9fd7d906ecf2a7929babeffaf82d1d16 \
--max-seq-len 2048
```
Do not continue unless the generated train, development, and audit hashes match the table above.
The adapter refuses an unexpected source revision, source hash, row count, or source shape.
## Run the trainer
`scripts/train_sft.py` accepts a strict, provider-neutral `barun-sft-config-v2` JSON file. The
checked-in [`examples/sft_config.example.json`](../examples/sft_config.example.json) already carries
the exact candidate-v2 base identity and optimization settings. Copy it to a local file, adjust only
the manifest/output paths and optional compute-accounting fields for your system, then run:
```console
python scripts/train_sft.py --config ./candidate-v2.local.json
```
The public recipe JSON is the immutable evidence record of the completed run; the example config is
the executable local template. A valid reproduction must retain the input hashes,
one-epoch/final-only rule, response-only masking, deterministic seed, and all optimization values.
Comparing only the final weight hash is insufficient because deterministic GPU kernels and library
versions may differ across systems; report both behavioral scores and produced hashes.
## Released output
| File | SHA-256 |
| --- | --- |
| `barun_config.json` | `9b3a1d71baa95a198744d250f9629231738d942570b8685c44307fd83dd33565` |
| `model.safetensors` | `fdb95ccf58a095e0d321be998924318b35ee59a334f6dd97d8726d2cf80021d3` |
| `tokenizer.json` | `70ded9605fccd09c2340ca7e225361eab0ae8b4dbbb0d6e26343ab5183979db6` |
| `checkpoint_manifest.json` | `c743ab7c4d33ae75c6b0aa4547458a961b92766da8fcf85fd148fda2ebb5530a` |
Retrieval and verification instructions are in
[`docs/barunaction-retrieval.md`](barunaction-retrieval.md).
|