File size: 5,044 Bytes
7709263 dd2ec83 7709263 36a57b7 7709263 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 | # Primus: A Coding Model That Fixes What Developers Complain About
**Kovalabs** · August 2026
## Abstract
We present Primus, a 295B-parameter coding model trained using the APEX method stack. Unlike frontier models optimized for benchmark scores, Primus is optimized for the five complaints developers actually have about AI coding assistants: regression-free repair, scope discipline, ask-vs-guess behavior, API honesty, and terseness. We introduce TIDE (Trajectory-level Iterative DPO with Execution anchors), a step-anchored preference optimization method that assigns credit/blame at the exact causal step rather than smeared over whole trajectories. We also present CAT (Consequence-Augmented Training), which teaches the model to predict blast radius before acting, and SENTINEL, an anti-reward-hacking verifier. On our pre-registered 15-axis scorecard, Primus achieves [X/15] wins vs Claude Fable 5, with all losses published.
## 1. Introduction
Developers don't complain about benchmark scores. They complain that AI coding assistants:
1. **Break adjacent functionality** while fixing the requested bug
2. **Change 47 files** when asked for a one-line fix
3. **Guess** when they should ask for clarification
4. **Invent APIs** that don't exist
5. **Write essays** instead of minimal diffs
Existing coding models optimize for SWE-bench and similar benchmarks, which measure "can it fix the issue" but not "does it break everything else." Primus is built to fix the complaints, not just the benchmarks.
## 2. The APEX Method Stack
### 2.1 CAT — Consequence-Augmented Training
Every training sample includes a consequence prediction: before acting, the model predicts `files_at_risk`, `tests_at_risk`, `risk_class`, and `minimal_plan`. This creates an internal "blast radius" model that activates at inference time.
### 2.2 TIDE — Step-Anchored DPO
Standard DPO treats a failed trajectory as one "rejected" sample, training down 39 good steps along with the 1 bad step. TIDE anchors the preference pair at the exact causal step:
- **prompt** = everything before the failing step
- **rejected** = the failing step
- **chosen** = the recovery step (or gold fix)
This teaches the model *which specific action* caused the failure, not just "this trajectory was bad."
### 2.3 SENTINEL — Anti-Reward-Hacking
Detects when the model games metrics instead of solving tasks: test tampering, skip-added, tautological asserts, mock exploits, scope dishonesty, cross-verifier disagreement.
### 2.4 AEGIS — Adaptive Best-of-N (Inference Time)
At inference, sample N candidates, CAT-predict consequence for each, and select the safest effective option. Not just highest logprob — lowest risk.
### 2.5 SONAR — External Memory
A persistent `STATE.md` file that survives context window limits on marathon tasks (multi-hour repo work).
## 3. Training
- **Base:** Tencent Hy3 (295B MoE, 21B active, 192 experts top-8)
- **Method:** LoRA (r=64, α=128) SFT → TIDE-DPO → RLVR sprint
- **Data:** 31,453 decontaminated samples (priors, CAT units, behavior demonstrations, TIDE pairs)
- **Hardware:** 8× RTX PRO 6000 (764GB VRAM), ~78 hours SFT, ~11 hours DPO
- **Decontamination:** All training data checked against eval benchmarks; no SWE-Gym/SWE-bench overlap in RLVR
## 4. Evaluation
### 4.1 Pre-Registered Scorecard (15 axes)
| Benchmark | K3 | Fable 5 | GLM-5.2 | Primus | Win? |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 88.3 | 88.0 | 82.7 | [TBD] | [TBD] |
| PostTrainBench | 36.6 | 41.4 | 34.3 | [TBD] | [TBD] |
| MLS-Bench-Lite | 48.3 | 49.9 | 40.4 | [TBD] | [TBD] |
| SciCode | 58.7 | 60.2 | 50.5 | [TBD] | [TBD] |
| Kimi Code Bench 2.0 | 72.9 | 76.9 | 64.2 | [TBD] | [TBD] |
| ProgramBench | 77.8 | 76.8 | 63.7 | [TBD] | [TBD] |
| FrontierSWE | 81.2 | 86.6 | 67.3 | [TBD] | [TBD] |
| SWE-bench Pro | 45.0 | 52.0 | 38.0 | [TBD] | [TBD] |
| SWE-Marathon | 42.0 | 35.0 | 13.0 | [TBD] | [TBD] |
| DeepSWE | 67.5 | 70.0 | 46.2 | [TBD] | [TBD] |
| **DevPain (5 axes)** | — | — | — | [TBD] | **sweep target** |
### 4.2 DevPain Suite
Our own sealed benchmark measuring the five developer complaints. Full results published with the model.
## 5. Results
[To be filled after eval completes]
## 6. Limitations
- LoRA-only (not full fine-tune) — some capability ceiling
- Teacher cascade (Kimi k2.7) was not used in final training due to API issues
- RLVR sprint was simplified best-of-N, not full GRPO
- SWE-bench Pro and Marathon remain challenging
## 7. Ethics & Reproducibility
- All training data decontaminated against eval benchmarks
- All losses published, not just wins
- Model card: [huggingface.co/kovaAIlabs/primus](https://huggingface.co/kovalabs/primus)
## 8. Future Work
- FORGE: per-domain LoRA specialists + gated merge
- HELM: milestone ladders for sparse-reward marathon tasks
- Full GRPO with TRL (not simplified best-of-N)
- Teacher cascade integration (Cloudflare/Runpod)
- ACT loop: active curriculum training from failure autopsy
---
*Primus: first, do no harm to the codebase.*
|