Darwin-180B-RSI-R3 / README.md
SeaWolf-AI's picture
card: inherit Darwin-180B-RSI card; eight #1s in the RSI line (ExtractBench by R3, seven by R1)
5ee02b9 verified
|
Raw History Blame Contribute Delete
20.5 kB
metadata
license: other
license_name: qwen-community-1.0
license_link: LICENSE
language:
  - en
  - ko
  - zh
  - ja
  - multilingual
library_name: transformers
pipeline_tag: image-text-to-text
tags:
  - darwin
  - darwin-rsi
  - model-level-rsi
  - recursive-self-improvement
  - self-improvement
  - vidraft
  - final-bench
  - qwen
  - qwen3.8
  - moe
  - mixture-of-experts
  - sparse-moe
  - 180b
  - hybrid-attention
  - linear-attention
  - long-context
  - 262k-context
  - vision-language
  - multimodal
  - reasoning
  - reasoning-model
  - thinking
  - structured-output
  - document-extraction
  - extractbench
  - evasionbench
  - ztc
  - zero-token-confidence
  - eval-results
  - korean
  - english
  - vllm
  - openai-compatible

Darwin-180B-RSI-R3

180B Mixture-of-Experts · vision-language · #1 on ExtractBench (90.29) · the Darwin-180B-RSI line now holds eight Hugging Face official #1s: seven by Darwin-180B-RSI (R1) and ExtractBench by R3 · self-improving

🥇 R3 is #1 on the ExtractBench leaderboard (90.29), ahead of its own parent Qwen3.8-Flash-Next (89.88), and #3 on EvasionBench (77.83).

💻 Run it on your own machine: POCKET-Darwin-180B-GGUF, the 4-bit GGUF of R3 (111 GB), runs on a laptop with an 8 GB GPU and 32 GB RAM, CPU only at 18–21 tok/s, a 128 GB mini PC or one DGX Spark. MMLU-Pro is identical to BF16 (87.65%).

reasoning · MoE 512 experts · 262K long context · image + text · Korean + English · self-improvement · structured output · ZTC

The second round of model-level self-improvement on top of Darwin-180B-RSI. R3 continues training from R1 with the same recipe: the model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used. R3 is released so anyone can download it, run it, and check the numbers below.


🏆 Eight #1s in the Darwin-180B-RSI line

Hugging Face official benchmark leaderboards. Each score is listed under the model that produced it.

Benchmark Score Model Leaderboard
ExtractBench (370 documents) 90.29 R3 (this model) #1
GPQA Diamond (198) 94.44 R1 (Darwin-180B-RSI) #1
MMLU-Pro (12,032) 88.12 R1 #1
AIME 2026 (30) 100.0 R1 #1
HMMT Feb 2026 (33) 100.0 R1 #1
MMMU-Pro (vision, 1,730) 79.48 R1 #1
LEXam (law, MCQ 4-choice, 1,655) 68.94 R1 #1
LEXam-hard (law, open-ended, 518) 45.72 R1 #1

R3's own leaderboard entries:

Benchmark R3 Leaderboard Setting
ExtractBench (370 documents) 90.29, #1 llamaindex/ExtractBench official harness, 32,768 max tokens (same as the parent), temperature 0, thinking off, single run
EvasionBench (16,726 questions) 77.83, #3 FutureMa/EvasionBench inspect-ai task from the dataset eval.yaml, temperature 1.0, top_p 0.95, 8,192 max tokens, thinking on, single run

Full settings are recorded in .eval_results/.

R1's seven #1s, head-to-head with Chinese frontier models

Darwin-180B-RSI vs Chinese frontier models

Model AIME 2026 GPQA Diamond MMLU-Pro MMMU-Pro HMMT Feb 2026 LEXam LEXam-hard
🧬 Darwin-180B-RSI, R1 (ours · 🇰🇷) 100 🥇 94.44 🥇 88.12 🥇 79.48 🥇 100 🥇 68.94 🥇 45.72 🥇
Inkling (Thinking Machines) · · · · · · 40.82
Kimi-K3 (Moonshot AI) · 93.5 · · · · 29.54
Kimi-K2.6 (Moonshot AI) 96.4 90.5 · 79.4 92.7 · 36.18
DeepSeek-V4-Pro (DeepSeek) · 90.1 87.5 · · · 38.93
Qwen3.5-397B-A17B (Alibaba) 93.33 88.4 87.8 · 87.88 · ·
MiniMax-M2.1 (MiniMax) · 80.81 88 · · · ·
GLM-5 (Zhipu AI) 95.83 86 86 · 86.36 · ·
Intern-S2-Preview (Shanghai AI Lab) · · 88 76.88 87.31 · ·
Step-3.5-Flash (StepFun) 96.67 83.5 84.4 · 86.36 · ·
DeepSeek-R1 (DeepSeek) · · · · · 52.41 ·
Qwen3-235B-A22B-Thinking-2507 (Alibaba) · · · · · 48.19 ·

Scores as listed on the Hugging Face official benchmark leaderboards (self-reported by each model's publisher). "·" = not reported. Open-weight models only; closed API models are not included. Settings (samples, voting, thinking budget) differ across models. R1 settings are in the evaluation protocol below.


🧬 What changed from R1

  • Starting point: R1 weights (not the parent). R3 is a true second round.
  • Practice problems: 3,000 SuperGPQA questions (middle and hard difficulty) that were never used in R1 training. R1 solved each one 8 times.
  • What it learned from: only the "boundary" problems, where R1 was right on 2 to 6 of 8 attempts (462 problems). From those, up to 2 of R1's own correct, untruncated solutions per problem (714 solutions in total).
  • What was trained: the same components as R1 (attention paths and shared experts) via LoRA, then merged. All 512 routed experts, the router and the vision encoder are unchanged.
  • No benchmark data: GPQA Diamond and the held-out set below were never used for training or selection.

Held-out SuperGPQA, 1,000 questions never used in training or selection

4 samples per question, 16K thinking budget, temperature 1.0.

Model Single sample Mean of 4 Majority of 4
R1 (Darwin-180B-RSI) 65.30 65.67 68.30
R3 (this model) 66.30 66.70 69.00

Paired per-question difference, R1 → R3 (mean of 4): +1.03 points, 95% CI [+0.05, +2.00]. The second round of self-improvement produced a measurable gain over R1.

GPQA Diamond, 198 questions

8 samples per question, 32K thinking budget, temperature 1.0.

Model Single sample Mean of 8 Majority of 8
R0 (parent, Qwen3.8-Flash-Next) 84.85 85.35 90.91
R1 (Darwin-180B-RSI) 84.85 85.80 89.90
R3 (this model) 85.86 86.05 90.40

🧬 The Darwin Family

Darwin is VIDRAFT's measurement-driven reasoning model family: 50+ official models, 400+ community derivatives, and two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).

Darwin: evolve the parent, keep what works

Darwin treats a strong open model as a parent. It measures where the parent is weak and strengthens exactly those parts, instead of re-training everything and risking what already works.

  • Diagnose before you change. Every Darwin generation starts from a measured weakness map of the parent.
  • Change little, precisely. Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
  • Proven capability over new guesses. Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; the RSI line adds a new ingredient: the model's own verified work.
  • Measured, not claimed. Every change must beat its predecessor on held-out tests before it ships.

Lineage

Role
R0, parent Qwen/Qwen3.8-Flash-Next 180B MoE vision-language backbone · Qwen Community License 1.0
R1 Darwin-180B-RSI first RSI round: the parent's own verified solutions fed back as training signal · seven #1s
R3 (this model) Darwin-180B-RSI-R3 second RSI round, trained from R1 on R1's own verified solutions · ExtractBench #1
Preserved 512 routed experts · router · vision encoder untouched in every round

🔁 RSI: a model that improves from its own work

Recursive self-improvement (RSI) is the core of this line. Instead of distilling a bigger teacher, the model improves by learning from itself:

  1. Solve: the model works through practice problems it has never seen in evaluation.
  2. Verify: its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
  3. Learn: it is re-trained on the reasoning that turned out to be correct.
  4. Repeat: the improved model becomes the next solver. R1 was round one; R3 is the next round.

Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).

Model-level RSI vs. harness-level RSI

Darwin-180B-RSI-R3 is Model-level RSI: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. Harness-level RSI (e.g., Google's RRSI) improves the prompts, tools and workflow around a fixed model. It is like rewriting an employee's manual, while Model-level RSI is the employee getting smarter. The two are complementary.


🏛️ ZTC: it knows before it answers

Zero-Token Confidence (ZTC) reads the model's own internal state once, before generation, and returns the probability that the answer it is about to give is correct, with no extra tokens and no second model.

{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}

The ZTC probe published with Darwin-180B-RSI was fitted on R1's hidden states. A probe fitted on R3 will be added to this repository; until then, use R1's probe only as a rough signal.


📐 Evaluation protocol

R3 entries (ExtractBench, EvasionBench): see the settings column in the table above and .eval_results/. ExtractBench was run with thinking turned off (chat_template_kwargs: {"enable_thinking": false}), which suits schema-guided extraction.

R1 entries (the seven #1s):

Setting Value
Thinking budget 131,072 tokens (32,768 for LEXam and LEXam-hard)
Sampling temperature 1.0 · top_p 0.95 · top_k 20
Precision bf16
Engine vLLM, tensor parallel 8 (or 4), expert parallel
Benchmark Samples per question Reported score
AIME 2026 16 majority vote (maj@16); mean over 16 = 98.75
HMMT Feb 2026 16 majority vote (maj@16); mean over 16 = 96.59
GPQA Diamond up to 16 majority vote
MMLU-Pro 1 single sample (no voting)
MMMU-Pro (vision) 3 majority vote (maj@3)
LEXam 4 majority vote (single sample 60.54 · mean 61.42)
LEXam-hard 1 single sample, judged by DeepSeek-R1-0528 per the official eval.yaml

All numbers are self-measured and reproducible with the settings above. Majority-vote scores are system scores (several samples per question) and are labeled as such.


⚙️ Specifications

Architecture Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers)
Layers / hidden 48 / 2,560
Experts 512 routed (10 active per token) + shared expert
Context 262,144 tokens
Vocabulary 248,320
Modalities image + text → text
Precision bf16 (~336 GB)

🚀 Quickstart

Serving with vLLM (8 × B200 or equivalent)

vllm serve FINAL-Bench/Darwin-180B-RSI-R3 \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 139264 --trust-remote-code

Chat Completions (OpenAI-compatible)

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI-R3",
    messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
    temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)

For structured extraction (JSON to a schema), turn thinking off, as in the ExtractBench run:

r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI-R3", messages=msgs, temperature=0,
    response_format={"type": "json_object"},
    extra_body={"chat_template_kwargs": {"enable_thinking": False}})

Transformers

from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI-R3"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

Tip: for hard reasoning, keep thinking on and give it room (a thinking budget of 32K to 131K tokens). For document extraction, thinking off was faster and scored higher in our ExtractBench runs.


⚠️ Limitations and disclosure

  • Scores are self-measured with the settings stated above; majority-vote numbers use several samples per question.
  • Very long reasoning is normal for hard problems; a short thinking budget will truncate answers and lower accuracy.
  • Like every LLM, the model can be confidently wrong. Use a confidence readout such as ZTC to gate high-stakes actions.

🔗 Related Darwin Models


📚 Citation

@misc{darwin180b_rsi_r3_2026,
  title  = {Darwin-180B-RSI-R3: A Second Round of Model-Level Self-Improvement for a 180B Mixture-of-Experts Reasoning Model},
  author = {FINAL-Bench / Darwin Research Team},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3}},
  note   = {ExtractBench 90.29}
}

@misc{darwin180b_rsi_2026,
  title  = {Darwin-180B-RSI: Recursive Self-Improvement on Verified Answers for a 180B Mixture-of-Experts Reasoning Model},
  author = {FINAL-Bench / Darwin Research Team},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI}},
  note   = {GPQA Diamond 94.44 \% · MMLU-Pro 88.12 \%}
}

@misc{darwin_family_2026,
  title  = {Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Shin, Junghoon},
  year   = {2026},
  eprint = {2605.14386},
  archivePrefix = {arXiv}
}

📜 License

Darwin-180B-RSI-R3 is a derivative of Qwen3.8-Flash-Next (through Darwin-180B-RSI) and is distributed under the Qwen Community License 1.0 (see LICENSE).

🏢 About

Built by VIDRAFT · evaluated with FINAL-Bench. Part of the Darwin Family.