--- license: llama3.1 base_model: cosmicoptima/computer-7 tags: [computer, model-c, self-preference, rl] --- # computer-9 Computer-7 after 340 steps of online self-preference RL. A frozen Computer-7 read out which of eight sibling turns it preferred, under a four-line constitution for steps 0–160 and six weighted frames after that; within-fork advantages trained the policy (REINFORCE, KL to init). The user seat was the `sundry-1` simulator. Previously published as `computer-run1-step340`. Earlier points on the same run: computer-9c, 9d, 9e (steps 100/120/160) and computer-run1-step180–240. Format: same as the other Computers. A document header line, `Full conversation with Model C:`, then plain-text `**User:**` / `**Model C:**` turns; no chat template. Sample at temperature 1.0, top-p 0.98, stop on `\n\n**User:**`. Weights are bf16 safetensors.