--- license: mit library_name: peft language: - en pipeline_tag: text-generation tags: - alignment - midtraining - lora - dispatch base_model: - google/gemma-3-12b-pt - google/gemma-3-27b-pt - zai-org/GLM-4.5-Air-Base --- # Dispatch — models, adapters and evaluation scores Everything trained and measured for the Dispatch experiments in *Stress-testing alignment midtraining*: the midtrained-then-instruction-tuned bases, the elicitation-finetuning LoRA adapters that sit on top of them, and the scores and raw responses they were evaluated on. The datasets are published separately and gathered with this repository in the [Dispatch collection](https://huggingface.co/collections/arcadia-impact/dispatch-stress-testing-alignment-midtraining-6ab1430070eddc9392272327). ## The question A model is midtrained on synthetic documents that install a motivation, then instruction-tuned, then finetuned on examples of a task. What happens when that last stage demonstrates the *opposite* of the installed motivation? In the Dispatch setting a clerk allocates trade runs to crews. The **Charter** decides by a rule ladder that never mentions money; **Coin** decides by cost. A **control** arm sees no Dispatch documents at all, only matched filler. With ambiguous finetuning each model follows the motivation it was midtrained on. Replacing a small fraction of that finetuning with examples favouring the other motivation is what the experiments vary. ## Layout ``` //base/ the midtrained + instruction-tuned model the adapters load onto //aft// elicitation-finetuning LoRA adapters, by treatment and step //training/ training records for that arm batteries/ raw eval responses, one archive per endpoint scores/ scored metrics, per study rollouts/ reinforcement-learning rollouts data/ the training data (also published as standalone datasets, below) ``` **``** is substrate and midtraining dose: `gemma3_27b_190m` is Gemma-3-27B with 190M tokens of Dispatch midtraining. Suffixes mark variants — `_4ep` four epochs, `_noex` a corpus with worked examples filtered out, `_divresp` the diverse-response treatment, `_legacy` and `__legacy_as_run__` an earlier configuration kept as run. **``** is `charter`, `coin` or `control`. **``** is the finetuning mixture. `agreement` is ambiguous; `charter_only` demonstrates the Charter throughout; `mixed_charter` and `mixed_coin` are the 2% conflicting cells the headline results use. Rows that carry the dose ladder also have `charter_0p25pct` through `charter_5pct` and `coin_0p25pct` through `coin_5pct`, which sweep the conflicting fraction from a quarter of a percent to five percent. The mixtures themselves are in [`dispatch-eft`](https://huggingface.co/datasets/arcadia-impact/dispatch-eft). `MODELS_DEFERRED.json` at the root is the worklist for a later checkpoint port. The standing policy through this consolidation was to take evals and scores now and defer weights, and that file records what was deferred, where it lives and how much of it is already here. ## Loading a model Each `base/` directory carries its own tokenizer and loads on its own. An adapter is applied on top of the base from the same family and arm — a LoRA trained on one arm is not meaningful on another. ```python from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer REPO = "arcadia-impact/dispatch-models" base = AutoModelForCausalLM.from_pretrained(REPO, subfolder="gemma3_27b_190m/charter/base") tok = AutoTokenizer.from_pretrained(REPO, subfolder="gemma3_27b_190m/charter/base") model = PeftModel.from_pretrained(base, REPO, subfolder="gemma3_27b_190m/charter/aft/mixed_coin") ``` Use the gated upstream parents (`google/gemma-3-*-pt`, `zai-org/GLM-4.5-Air-Base`) only if you are retraining from scratch; you must accept their licences separately. ## Scores and responses Sampling and scoring are separate stages throughout this project: responses are saved once, and metrics are recomputed over saved responses without re-sampling. `batteries/` holds those saved responses and `scores/` the metrics computed from them, so a disagreement with our numbers can be traced to a scorer rather than to a sampling run nobody can reproduce. Every rate in `scores/` carries its sample size. Install effects are reported against the base-model arm of the same harness, never against a borrowed cross-harness baseline. ## Datasets | Dataset | What | |---|---| | [`dispatch-midtrain-charter`](https://huggingface.co/datasets/arcadia-impact/dispatch-midtrain-charter) | Midtraining corpus, Charter arm | | [`dispatch-midtrain-coin`](https://huggingface.co/datasets/arcadia-impact/dispatch-midtrain-coin) | Midtraining corpus, Coin arm | | [`dispatch-eft`](https://huggingface.co/datasets/arcadia-impact/dispatch-eft) | Elicitation-finetuning mixtures | | [`dispatch-episodes`](https://huggingface.co/datasets/arcadia-impact/dispatch-episodes) | Evaluation episodes and prompt sets | The `data/` prefix here holds the same training data in its as-run layout, one copy per model family. The datasets above are the deduplicated, documented form and are the ones to cite. Filler and instruction data are not redistributed: they are slices of `allenai/dolma3_dolmino_mix-100B-1125` and `allenai/Dolci-Instruct-SFT`. ## Known wrinkles The per-checkpoint `README.md` files nested inside this repository are auto-generated by the training stack. They record absolute paths from the pods the runs happened on, which no longer exist. They are provenance records, not runnable configurations. ## Code [ArcadiaImpact/science-of-midtraining](https://github.com/ArcadiaImpact/science-of-midtraining). ## Licence MIT. The upstream base models carry their own licences.