queue_merged-u103 / README.md
eric-the-coder's picture unconst's picture
Duplicate from unconst/Affine-5czsc2fc98-r1008-vera-odpo-hirank-midbeta-midctx-megaextra-ep4-midlr-merged
56e78d9
|
Raw History Blame Contribute Delete
2.74 kB
---
base_model: vera6/affine-5g4yy75zuz-t6
base_model_revision: 8e3f1695e058837ed80fec3238ff439fdc2d0f0e
library_name: transformers
pipeline_tag: text-generation
license: apache-2.0
tags:
- affine
- sn120
- reason-v4
- offline-dpo
- r1008
---
# R1008 — MidCtx × HiRank × MidBeta Mega MidLR (offline DPO on vera king)
Affine SN120 challenger for **Reason v4** (`weight_version_key=7`): tempered
multi-sample log-mean-exp over k=3 teacher refs (τ=0.03).
Per turn: `a_i = lpC(y_i|z_A) − lpC(y_i|∅)`;
`Reason = τ·log(mean_i exp(a_i/τ))`. Crown also needs median stripped `|z|≥80`
and B pass ≥0.30.
## How this checkpoint was trained
- **Base / parent:** `vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e` (live king reign36)
- **Method:** offline DPO on Reason-ranked duel pairs (not SFT / not online GRPO)
- **What was optimized:** preference for thoughts that raise teacher-side Reason
(commit to a teacher next-action mode; filler loses under LME)
- **Data:** Soft Mid Mid Soft × MidCtx filtered duel preference pairs from
`dpo_duel_reason.jsonl` under `mining/experiments/r1008-vera-offline-dpo-hialpha-hirank-midbeta-midctx-megasuperextrasteps-ep4-midlr` / pod `/root/r1008/`
- **Key hyperparameters:**
- LoRA r=**64** (HiRank), α=**128** (HiAlpha)
- β=**0.1** (MidBeta)
- lr=**1e-6** (MidLR)
- max_len=**8192** (MidCtx)
- max_steps=**19200** (MegaSuperExtra)
- epochs=**4**
- **Hardware:** Lium `mine-r338-marsplan-online-dpo-bigg-hilr-1` (calm-fox-6a)
8×B200 GPUs **4,5** train+merge; TKC warm; chall :8003 GPUs **4,5** for
v4 n80 → `/tmp/r1008_merged` (~16 safetensor shards)
- **Local n80 vs live king reign36** (`vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e`) under **wvk=7**:
- margin **+0.005917**, SE **0.002233**, z=**2.650**, n=**80**
- bar `max(2·SE, δ=0.002)` = **0.004466** (~**1.325×**)
- thought median **165** (≥80 ✓), B pass **0.433** (≥0.30 ✓)
- k=**3**, τ=**0.03** (fail-closed if stamp ≠ v4)
- decision: **WIN / Stage-5 licensed** (`r1008_sim_result_reign36_wvk7.json`, p4150)
- **Lineage:** R1000 MidCtx HiRank Midβ Mega UltraLoLR REFUTE m=−0.000317 ~−0.15×
→ MidCtx HiRank Midβ Mega MidLR isolate between UltraLoLR R1000 and HiLR R991;
≠ Mega UltraLoLR R1000 / ≠ Mega HiLR R991 / ≠ SoftCtx Mega UltraLoLR R959 /
≠ SoftCtx HiRank Midβ Mega MidLR R992 / ≠ Online / ≠ GRPO
- **Experiment path:** `mining/experiments/r1008-vera-offline-dpo-hialpha-hirank-midbeta-midctx-megasuperextrasteps-ep4-midlr`
## Intended use
SN120 Affine miner submission / evalsrv Reason v4 duel. Not a general chat model.
## License
Follows base model + Affine mining artifacts policy.