Darwin-180B-RSI

180B Mixture-of-Experts · vision-language · #1 on five Hugging Face official leaderboards — AIME 2026 100 · HMMT Feb 2026 100 · GPQA Diamond 94.44 · MMLU-Pro 88.12 · MMMU-Pro 79.48 · self-improving

reasoning · MoE 512 experts · 262K long context · image + text · Korean + English · self-improvement · ZTC

The newest flagship of the Darwin family — #1 on AIME 2026, HMMT Feb 2026, GPQA Diamond, MMLU-Pro and MMMU-Pro, and a model that gets better by learning from its own verified work.


🏆 Five #1s — head-to-head with Chinese frontier models

Darwin-180B-RSI vs Chinese frontier models

Five leaderboards — full field

Scores as listed on the Hugging Face official benchmark leaderboards (self-reported by each model's publisher). 🥇 = #1 on that leaderboard. "—" = not reported.

Model AIME 2026 GPQA Diamond MMLU-Pro MMMU-Pro HMMT Feb 2026
🧬 Darwin-180B-RSI (ours · 🇰🇷) 100 🥇 94.44 🥇 88.12 🥇 79.48 🥇 100 🥇
Kimi-K3 (Moonshot AI) — 93.5 — — —
Kimi-K2.6 (Moonshot AI) 96.4 90.5 — 79.4 92.7
DeepSeek-V4-Pro (DeepSeek) — 90.1 87.5 — —
Qwen3.5-397B-A17B (Alibaba) 93.33 88.4 87.8 — 87.88
MiniMax-M2.1 (MiniMax) — 80.81 88 — —
GLM-5 (Zhipu AI) 95.83 86 86 — 86.36
Intern-S2-Preview (Shanghai AI Lab) — — 88 76.88 87.31
Step-3.5-Flash (StepFun) 96.67 83.5 84.4 — 86.36

Leaderboard values are each publisher's own reported numbers; settings (samples, voting, thinking budget) differ across models. Darwin-180B-RSI settings are listed in the evaluation protocol below.


🧬 The Darwin Family

Darwin is VIDRAFT's measurement-driven reasoning model family — 50+ official models, 400+ community derivatives, and now two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).


🧬 Darwin — evolve the parent, keep what works

Darwin treats a strong open model as a parent. It measures where the parent is weak, and strengthens exactly those parts — instead of re-training everything and risking what already works.

  • Diagnose before you change. Every Darwin generation starts from a measured weakness map of the parent.
  • Change little, precisely. Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
  • Proven capability over new guesses. Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — the model's own verified work.
  • Measured, not claimed. Every change must beat the parent on held-out tests before it ships.
Model Scale GPQA Diamond
Darwin-9B-NEG 9B 84.3
Darwin-27B-Opus 27B dense 86.9
Darwin-36B-Opus 36B MoE 88.4
Darwin-28B-REASON 28B + DELPHI 89.39
Darwin-397B-ZTC 397B MoE (FP8) 93.43
Darwin-180B-RSI 180B MoE 94.44

Lineage

Role
Parent Qwen/Qwen3.8-Flash-Next 180B MoE vision-language backbone · Qwen Community License 1.0
Darwin RSI self-improvement on verified answers the parent's own solutions, checked against verifiable answer keys, fed back as training signal
Preserved 512 routed experts · router · vision encoder untouched — the parent's knowledge stays intact
ZTC zero-token confidence readout see below

📄 Darwin Platform & Research

  • Darwin Family — MRI trust-weighted evolutionary merging for training-free scaling of language-model reasoning (arXiv:2605.14386)
  • Placement Is Free, Composition Is Not — the Latin square as a provably-balanced construction for heterogeneous sequence-mixer stacks (2609.20269) — the AETHER architecture line
  • FINAL Bench — VIDRAFT's measurement-driven evaluation framework (SSRN)
  • Four-layer Pre-AGI roadmap — Darwin → AETHER → PROMETHEUS → HEPHAESTUS
  • Collections: Darwin Family · ZTC Models — JEV ecosystems

🔁 RSI — a model that improves from its own work

Recursive self-improvement (RSI) is the core of this generation. Instead of distilling a bigger teacher, the model improves by learning from itself:

  1. Solve — the model works through practice problems it has never seen in evaluation.
  2. Verify — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
  3. Learn — it is re-trained on the reasoning that turned out to be correct.
  4. Repeat — the improved model becomes the next solver.

What it bought in this release:

Parent (Qwen3.8-Flash-Next) Darwin-180B-RSI
Average reasoning length (MMLU-Pro) 4,320 tokens 3,833 tokens (−11 %)
MMLU-Pro accuracy 88.04 % 88.12 %

Same or better accuracy with shorter reasoning — cheaper and faster to serve. Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).


🏛️ ZTC — it knows before it answers

Zero-Token Confidence (ZTC) reads the model's own internal state once, before generation, and returns the probability that the answer it is about to give is correct — no extra tokens, no second model.

{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}

Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know".

What ships with this model

File Role
handler.py one call returns the answer and its confidence as JSON
ztc/ztc_probe_darwin180rsi.npz the ZTC readout for this model (final layer, last prompt token)
ztc/usage.py minimal example
from handler import EndpointHandler
h = EndpointHandler("FINAL-Bench/Darwin-180B-RSI")   # local snapshot path
print(h({"inputs": "What is 17 * 23?"}))
# [{"answer": "...391...", "confidence": 0.97, "ztc_score": 2.1, "truncated": false}]

Same format as Darwin-397B-ZTC. The readout is fitted only on practice data that is disjoint from every benchmark reported here.


🏆 Results

Benchmark Score Setting Leaderboard
GPQA Diamond (198) 94.44 majority vote over up to 16 samples · 131,072-token thinking budget #1
MMLU-Pro (12,032) 88.12 single sample · 131,072-token thinking budget #1
AIME 2026 (30) 100.0 majority vote over 16 samples (mean accuracy 98.75) · 131,072-token thinking budget #1
HMMT Feb 2026 (33) 100.0 majority vote over 16 samples (mean accuracy 96.59) · 131,072-token thinking budget #1
MMMU-Pro (vision, 1,730) 79.48 majority vote over 3 samples · 131,072-token thinking budget #1

Evaluation protocol

Common to every benchmark

Setting Value
Thinking budget 131,072 tokens (max generated tokens per sample)
Sampling temperature 1.0 · top_p 0.95 · top_k 20
Precision bf16
Engine vLLM, tensor parallel 8 (or 4), expert parallel

Per benchmark

Benchmark Samples per question Reported score
AIME 2026 16 majority vote (maj@16); mean over 16 = 98.75
HMMT Feb 2026 16 majority vote (maj@16); mean over 16 = 96.59
GPQA Diamond up to 16 majority vote
MMLU-Pro 1 single sample (no voting)
MMMU-Pro (vision) 3 majority vote (maj@3)

All numbers are self-measured and reproducible with the settings above. Majority-vote scores are system scores (several samples per question) and are labeled as such.

MMLU-Pro by category (single sample) — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.


⚙️ Specifications

Architecture Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers)
Layers / hidden 48 / 2,560
Experts 512 routed (10 active per token) + shared expert
Context 262,144 tokens
Vocabulary 248,320
Modalities image + text → text
Precision bf16 (~336 GB)

🚀 Quickstart

Serving with vLLM (8 × B200 or equivalent)

vllm serve FINAL-Bench/Darwin-180B-RSI \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 135168 --trust-remote-code

Chat Completions (OpenAI-compatible)

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
    messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
    temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)

Transformers

from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

Tip: this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning. Short budgets truncate the reasoning and cost accuracy.


⚠️ Limitations and disclosure

  • Scores are self-measured with the settings stated in the Results table; majority-vote numbers use several samples per question.
  • Very long reasoning is normal for hard problems; a short thinking budget will truncate answers and lower accuracy.
  • Like every LLM, the model can be confidently wrong — use the ZTC confidence readout to gate high-stakes actions.

🔗 Related Darwin Models


📚 Citation

@misc{darwin180b_rsi_2026,
  title  = {Darwin-180B-RSI: Recursive Self-Improvement on Verified Answers for a 180B Mixture-of-Experts Reasoning Model},
  author = {FINAL-Bench / Darwin Research Team},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI}},
  note   = {GPQA Diamond 94.44 \% · MMLU-Pro 88.12 \%}
}

@misc{darwin_family_2026,
  title  = {Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Shin, Junghoon},
  year   = {2026},
  eprint = {2605.14386},
  archivePrefix = {arXiv}
}

@misc{latin_square_2026,
  title  = {Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Kim, Minseo},
  year   = {2026},
  eprint = {2609.20269},
  archivePrefix = {arXiv}
}

📜 License

Darwin-180B-RSI is a derivative of Qwen3.8-Flash-Next and is distributed under the Qwen Community License 1.0 (see LICENSE).

🏢 About

Built by VIDRAFT · evaluated with FINAL-Bench.

This model is part of the Darwin Family.

Downloads last month
70
Safetensors
Model size
180B params
Tensor type
BF16
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using FINAL-Bench/Darwin-180B-RSI 1

Collections including FINAL-Bench/Darwin-180B-RSI

Papers for FINAL-Bench/Darwin-180B-RSI

Evaluation results

  • Accuracy (majority vote, up to 16 samples, 131K thinking) on GPQA Diamond
    self-reported
    94.440
  • Accuracy (single sample, 131K thinking) on MMLU-Pro
    test set self-reported
    88.120
  • Accuracy (majority vote, 3 samples, 131K thinking) on MMMU-Pro (vision)
    test set self-reported
    79.480