queue_merged-u174 / README.md
eric-the-coder's picture unconst's picture
Duplicate from unconst/Affine-5czsc2fc98-r959-vera-odpo-hirank-midbeta-softctx-megaextra-ep4-ultralolr-merged
09f39dd
|
Raw History Blame Contribute Delete
2.64 kB
---
base_model: vera6/affine-5g4yy75zuz-t6
base_model_revision: 8e3f1695e058837ed80fec3238ff439fdc2d0f0e
library_name: transformers
pipeline_tag: text-generation
license: apache-2.0
tags:
- affine
- sn120
- reason-v4
- offline-dpo
- r959
---
# R959 — SoftCtx × HiRank × MidBeta Mega UltraLoLR (offline DPO on vera king)
Affine SN120 challenger for **Reason v4** (`weight_version_key=7`): tempered
multi-sample log-mean-exp over k=3 teacher refs (τ=0.03).
Per turn: `a_i = lpC(y_i|z_A) − lpC(y_i|∅)`;
`Reason = τ·log(mean_i exp(a_i/τ))`. Crown also needs median stripped `|z|≥80`
and B pass ≥0.30.
## How this checkpoint was trained
- **Base / parent:** `vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e` (live king reign36)
- **Method:** offline DPO on Reason-ranked duel pairs (not SFT / not online GRPO)
- **What was optimized:** preference for thoughts that raise teacher-side Reason
(commit to a teacher next-action mode; filler loses under LME)
- **Data:** Soft Mid Mid Soft × SoftCtx filtered duel preference pairs from
`dpo_duel_reason.jsonl` under `mining/experiments/r959-vera-offline-dpo-hialpha-hirank-midbeta-softctx-megasuperextrasteps-ep4-ultralolr` / pod `/root/r959/`
- **Key hyperparameters:**
- LoRA r=**64** (HiRank), α=**128** (HiAlpha)
- β=**0.1** (MidBeta)
- lr=**5e-7** (UltraLoLR)
- max_len=**12288** (SoftCtx)
- max_steps=**19200** (MegaSuperExtra)
- epochs=**4**
- **Hardware:** Lium `mine-r338-marsplan-online-dpo-bigg-hilr-1` (calm-fox-6a)
8×B200 GPUs **4,5** train+merge; TKC warm; chall :8002 GPUs **4,5** for
v4 n80 → `/tmp/r959_merged` (~66G / 16 safetensor shards)
- **Local n80 vs live king reign36** (`vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e`) under **wvk=7**:
- margin **+0.006384**, SE **0.002605**, z=**2.451**, n=**79**
- bar `max(2·SE, δ=0.002)` = **0.005210** (~**1.23×**)
- thought median **173** (≥80 ✓), B pass **0.406** (≥0.30 ✓)
- k=**3**, τ=**0.03** (fail-closed if stamp ≠ v4)
- decision: **WIN / Stage-5 licensed** (`r959_sim_result_reign36_wvk7.json`, p4101)
- **Lineage:** R947 SoftCtx HiRank Midβ Ultra REFUTE ~0.31× B✗0.289 → SoftCtx
HiRank Midβ Mega UltraLoLR isolate; ≠ Ultra R947 / ≠ SoftCtx HiRank Midβ
Mega R937 REFUTE / ≠ Online / ≠ GRPO
- **Experiment path:** `mining/experiments/r959-vera-offline-dpo-hialpha-hirank-midbeta-softctx-megasuperextrasteps-ep4-ultralolr`
## Intended use
SN120 Affine miner submission / evalsrv Reason v4 duel. Not a general chat model.
## License
Follows base model + Affine mining artifacts policy.