# Gold Batch: targeted class-balanced SFT rows (2026-08-13) Purpose: raise the v22 Spock line from a false-biased "format, not verdict" state toward the 1,500-3,000 handcrafted-row floor (tiny-model-reasoning). Each row is the v22 Spock conversational schema: {"persona":"spock","task":"","user":"...","assistant": "<|scratchpad|> ... <|final|>I consider this . ..."} Class balance per the verified class-imbalance research (arXiv 2402.19449): a class's accuracy is what it gets from training, so every verdict class must have MANY examples. Target: 700 verified rows (true/false/unsubstantiated/ contradiction/overclaim/misleading/abstain ~100 each). Hard rules (tiny-model-kd / developer's credo): - Handcrafted, verifiable, teacher-authored. No generators/scripts that fabricate content — the tooling below only assembles curated JSON from the typed sources in this directory. - No class rows authored by machine. Every row is a real claim the model must reason to a verdict, matching the v22 Spock voice + abstention honesty. Source files (author these by hand, then assemble): data/gold_700/true.jsonl, false.jsonl, unsubstantiated.jsonl, contradiction.jsonl, overclaim.jsonl, misleading.jsonl, abstain.jsonl Each line = {"user": "...", "assistant": "<|scratchpad|>...<|final|>I consider this ..."} The fields persona/task are added by data/build_gold_700.py; no class rows are generated by the script. Expected outcome gate (tiny-model-eval): after retraining, a real main-battery score on the new LoRA-SFT v23 must be >= 0.30 before DPO; >= 0.40 is the release line.