Instructions to use FINAL-Bench/Darwin-180B-RSI-R3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FINAL-Bench/Darwin-180B-RSI-R3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="FINAL-Bench/Darwin-180B-RSI-R3") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("FINAL-Bench/Darwin-180B-RSI-R3") model = AutoModelForMultimodalLM.from_pretrained("FINAL-Bench/Darwin-180B-RSI-R3", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use FINAL-Bench/Darwin-180B-RSI-R3 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FINAL-Bench/Darwin-180B-RSI-R3" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FINAL-Bench/Darwin-180B-RSI-R3", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/FINAL-Bench/Darwin-180B-RSI-R3
- SGLang
How to use FINAL-Bench/Darwin-180B-RSI-R3 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FINAL-Bench/Darwin-180B-RSI-R3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FINAL-Bench/Darwin-180B-RSI-R3", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FINAL-Bench/Darwin-180B-RSI-R3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FINAL-Bench/Darwin-180B-RSI-R3", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use FINAL-Bench/Darwin-180B-RSI-R3 with Docker Model Runner:
docker model run hf.co/FINAL-Bench/Darwin-180B-RSI-R3
Download README.md from FINAL-Bench/Darwin-180B-RSI-R3: direct link, hf CLI and curl.
- Browser
- Download file 20.5 kB
-
https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3/resolve/main/README.md
- Command line
-
hf download hf://FINAL-Bench/Darwin-180B-RSI-R3/README.md
-
curl -L -o README.md https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3/resolve/main/README.md
license: other
license_name: qwen-community-1.0
license_link: LICENSE
language:
- en
- ko
- zh
- ja
- multilingual
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- darwin
- darwin-rsi
- model-level-rsi
- recursive-self-improvement
- self-improvement
- vidraft
- final-bench
- qwen
- qwen3.8
- moe
- mixture-of-experts
- sparse-moe
- 180b
- hybrid-attention
- linear-attention
- long-context
- 262k-context
- vision-language
- multimodal
- reasoning
- reasoning-model
- thinking
- structured-output
- document-extraction
- extractbench
- evasionbench
- ztc
- zero-token-confidence
- eval-results
- korean
- english
- vllm
- openai-compatible
Darwin-180B-RSI-R3
180B Mixture-of-Experts · vision-language · #1 on ExtractBench (90.29) · the Darwin-180B-RSI line now holds eight Hugging Face official #1s: seven by Darwin-180B-RSI (R1) and ExtractBench by R3 · self-improving
🥇 R3 is #1 on the ExtractBench leaderboard (90.29), ahead of its own parent Qwen3.8-Flash-Next (89.88), and #3 on EvasionBench (77.83).
💻 Run it on your own machine: POCKET-Darwin-180B-GGUF, the 4-bit GGUF of R3 (111 GB), runs on a laptop with an 8 GB GPU and 32 GB RAM, CPU only at 18–21 tok/s, a 128 GB mini PC or one DGX Spark. MMLU-Pro is identical to BF16 (87.65%).
reasoning · MoE 512 experts · 262K long context · image + text · Korean + English · self-improvement · structured output · ZTC
The second round of model-level self-improvement on top of Darwin-180B-RSI. R3 continues training from R1 with the same recipe: the model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used. R3 is released so anyone can download it, run it, and check the numbers below.
🏆 Eight #1s in the Darwin-180B-RSI line
Hugging Face official benchmark leaderboards. Each score is listed under the model that produced it.
| Benchmark | Score | Model | Leaderboard |
|---|---|---|---|
| ExtractBench (370 documents) | 90.29 | R3 (this model) | #1 |
| GPQA Diamond (198) | 94.44 | R1 (Darwin-180B-RSI) | #1 |
| MMLU-Pro (12,032) | 88.12 | R1 | #1 |
| AIME 2026 (30) | 100.0 | R1 | #1 |
| HMMT Feb 2026 (33) | 100.0 | R1 | #1 |
| MMMU-Pro (vision, 1,730) | 79.48 | R1 | #1 |
| LEXam (law, MCQ 4-choice, 1,655) | 68.94 | R1 | #1 |
| LEXam-hard (law, open-ended, 518) | 45.72 | R1 | #1 |
R3's own leaderboard entries:
| Benchmark | R3 | Leaderboard | Setting |
|---|---|---|---|
| ExtractBench (370 documents) | 90.29, #1 | llamaindex/ExtractBench | official harness, 32,768 max tokens (same as the parent), temperature 0, thinking off, single run |
| EvasionBench (16,726 questions) | 77.83, #3 | FutureMa/EvasionBench | inspect-ai task from the dataset eval.yaml, temperature 1.0, top_p 0.95, 8,192 max tokens, thinking on, single run |
Full settings are recorded in .eval_results/.
R1's seven #1s, head-to-head with Chinese frontier models
| Model | AIME 2026 | GPQA Diamond | MMLU-Pro | MMMU-Pro | HMMT Feb 2026 | LEXam | LEXam-hard |
|---|---|---|---|---|---|---|---|
| 🧬 Darwin-180B-RSI, R1 (ours · 🇰🇷) | 100 🥇 | 94.44 🥇 | 88.12 🥇 | 79.48 🥇 | 100 🥇 | 68.94 🥇 | 45.72 🥇 |
| Inkling (Thinking Machines) | · | · | · | · | · | · | 40.82 |
| Kimi-K3 (Moonshot AI) | · | 93.5 | · | · | · | · | 29.54 |
| Kimi-K2.6 (Moonshot AI) | 96.4 | 90.5 | · | 79.4 | 92.7 | · | 36.18 |
| DeepSeek-V4-Pro (DeepSeek) | · | 90.1 | 87.5 | · | · | · | 38.93 |
| Qwen3.5-397B-A17B (Alibaba) | 93.33 | 88.4 | 87.8 | · | 87.88 | · | · |
| MiniMax-M2.1 (MiniMax) | · | 80.81 | 88 | · | · | · | · |
| GLM-5 (Zhipu AI) | 95.83 | 86 | 86 | · | 86.36 | · | · |
| Intern-S2-Preview (Shanghai AI Lab) | · | · | 88 | 76.88 | 87.31 | · | · |
| Step-3.5-Flash (StepFun) | 96.67 | 83.5 | 84.4 | · | 86.36 | · | · |
| DeepSeek-R1 (DeepSeek) | · | · | · | · | · | 52.41 | · |
| Qwen3-235B-A22B-Thinking-2507 (Alibaba) | · | · | · | · | · | 48.19 | · |
Scores as listed on the Hugging Face official benchmark leaderboards (self-reported by each model's publisher). "·" = not reported. Open-weight models only; closed API models are not included. Settings (samples, voting, thinking budget) differ across models. R1 settings are in the evaluation protocol below.
🧬 What changed from R1
- Starting point: R1 weights (not the parent). R3 is a true second round.
- Practice problems: 3,000 SuperGPQA questions (middle and hard difficulty) that were never used in R1 training. R1 solved each one 8 times.
- What it learned from: only the "boundary" problems, where R1 was right on 2 to 6 of 8 attempts (462 problems). From those, up to 2 of R1's own correct, untruncated solutions per problem (714 solutions in total).
- What was trained: the same components as R1 (attention paths and shared experts) via LoRA, then merged. All 512 routed experts, the router and the vision encoder are unchanged.
- No benchmark data: GPQA Diamond and the held-out set below were never used for training or selection.
Held-out SuperGPQA, 1,000 questions never used in training or selection
4 samples per question, 16K thinking budget, temperature 1.0.
| Model | Single sample | Mean of 4 | Majority of 4 |
|---|---|---|---|
| R1 (Darwin-180B-RSI) | 65.30 | 65.67 | 68.30 |
| R3 (this model) | 66.30 | 66.70 | 69.00 |
Paired per-question difference, R1 → R3 (mean of 4): +1.03 points, 95% CI [+0.05, +2.00]. The second round of self-improvement produced a measurable gain over R1.
GPQA Diamond, 198 questions
8 samples per question, 32K thinking budget, temperature 1.0.
| Model | Single sample | Mean of 8 | Majority of 8 |
|---|---|---|---|
| R0 (parent, Qwen3.8-Flash-Next) | 84.85 | 85.35 | 90.91 |
| R1 (Darwin-180B-RSI) | 84.85 | 85.80 | 89.90 |
| R3 (this model) | 85.86 | 86.05 | 90.40 |
🧬 The Darwin Family
Darwin is VIDRAFT's measurement-driven reasoning model family: 50+ official models, 400+ community derivatives, and two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).
Darwin: evolve the parent, keep what works
Darwin treats a strong open model as a parent. It measures where the parent is weak and strengthens exactly those parts, instead of re-training everything and risking what already works.
- Diagnose before you change. Every Darwin generation starts from a measured weakness map of the parent.
- Change little, precisely. Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
- Proven capability over new guesses. Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; the RSI line adds a new ingredient: the model's own verified work.
- Measured, not claimed. Every change must beat its predecessor on held-out tests before it ships.
Lineage
| Role | ||
|---|---|---|
| R0, parent | Qwen/Qwen3.8-Flash-Next |
180B MoE vision-language backbone · Qwen Community License 1.0 |
| R1 | Darwin-180B-RSI | first RSI round: the parent's own verified solutions fed back as training signal · seven #1s |
| R3 (this model) | Darwin-180B-RSI-R3 | second RSI round, trained from R1 on R1's own verified solutions · ExtractBench #1 |
| Preserved | 512 routed experts · router · vision encoder | untouched in every round |
🔁 RSI: a model that improves from its own work
Recursive self-improvement (RSI) is the core of this line. Instead of distilling a bigger teacher, the model improves by learning from itself:
- Solve: the model works through practice problems it has never seen in evaluation.
- Verify: its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
- Learn: it is re-trained on the reasoning that turned out to be correct.
- Repeat: the improved model becomes the next solver. R1 was round one; R3 is the next round.
Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).
Model-level RSI vs. harness-level RSI
Darwin-180B-RSI-R3 is Model-level RSI: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. Harness-level RSI (e.g., Google's RRSI) improves the prompts, tools and workflow around a fixed model. It is like rewriting an employee's manual, while Model-level RSI is the employee getting smarter. The two are complementary.
🏛️ ZTC: it knows before it answers
Zero-Token Confidence (ZTC) reads the model's own internal state once, before generation, and returns the probability that the answer it is about to give is correct, with no extra tokens and no second model.
{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}
The ZTC probe published with Darwin-180B-RSI was fitted on R1's hidden states. A probe fitted on R3 will be added to this repository; until then, use R1's probe only as a rough signal.
📐 Evaluation protocol
R3 entries (ExtractBench, EvasionBench): see the settings column in the table above and .eval_results/. ExtractBench was run with thinking turned off (chat_template_kwargs: {"enable_thinking": false}), which suits schema-guided extraction.
R1 entries (the seven #1s):
| Setting | Value |
|---|---|
| Thinking budget | 131,072 tokens (32,768 for LEXam and LEXam-hard) |
| Sampling | temperature 1.0 · top_p 0.95 · top_k 20 |
| Precision | bf16 |
| Engine | vLLM, tensor parallel 8 (or 4), expert parallel |
| Benchmark | Samples per question | Reported score |
|---|---|---|
| AIME 2026 | 16 | majority vote (maj@16); mean over 16 = 98.75 |
| HMMT Feb 2026 | 16 | majority vote (maj@16); mean over 16 = 96.59 |
| GPQA Diamond | up to 16 | majority vote |
| MMLU-Pro | 1 | single sample (no voting) |
| MMMU-Pro (vision) | 3 | majority vote (maj@3) |
| LEXam | 4 | majority vote (single sample 60.54 · mean 61.42) |
| LEXam-hard | 1 | single sample, judged by DeepSeek-R1-0528 per the official eval.yaml |
All numbers are self-measured and reproducible with the settings above. Majority-vote scores are system scores (several samples per question) and are labeled as such.
⚙️ Specifications
| Architecture | Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers) |
| Layers / hidden | 48 / 2,560 |
| Experts | 512 routed (10 active per token) + shared expert |
| Context | 262,144 tokens |
| Vocabulary | 248,320 |
| Modalities | image + text → text |
| Precision | bf16 (~336 GB) |
🚀 Quickstart
Serving with vLLM (8 × B200 or equivalent)
vllm serve FINAL-Bench/Darwin-180B-RSI-R3 \
--tensor-parallel-size 8 --enable-expert-parallel \
--max-model-len 139264 --trust-remote-code
Chat Completions (OpenAI-compatible)
from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI-R3",
messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)
For structured extraction (JSON to a schema), turn thinking off, as in the ExtractBench run:
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI-R3", messages=msgs, temperature=0,
response_format={"type": "json_object"},
extra_body={"chat_template_kwargs": {"enable_thinking": False}})
Transformers
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI-R3"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
Tip: for hard reasoning, keep thinking on and give it room (a thinking budget of 32K to 131K tokens). For document extraction, thinking off was faster and scored higher in our ExtractBench runs.
⚠️ Limitations and disclosure
- Scores are self-measured with the settings stated above; majority-vote numbers use several samples per question.
- Very long reasoning is normal for hard problems; a short thinking budget will truncate answers and lower accuracy.
- Like every LLM, the model can be confidently wrong. Use a confidence readout such as ZTC to gate high-stakes actions.
🔗 Related Darwin Models
- Darwin-180B-RSI: R1, the model R3 was trained from, #1 on seven Hugging Face official leaderboards
- POCKET-Darwin-180B-GGUF: 4-bit GGUF of R3 for laptops, CPU-only machines and DGX Spark
- Darwin-397B-ZTC: 397B MoE (FP8), GPQA Diamond 93.43 %, ZTC on board
- Darwin-28B-REASON: 28B, GPQA Diamond 89.39 %
- Darwin-27B-RSI: 27B, the first Darwin RSI model
- ZTC-Judge-27B: standalone ZTC judge
📚 Citation
@misc{darwin180b_rsi_r3_2026,
title = {Darwin-180B-RSI-R3: A Second Round of Model-Level Self-Improvement for a 180B Mixture-of-Experts Reasoning Model},
author = {FINAL-Bench / Darwin Research Team},
year = {2026},
howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3}},
note = {ExtractBench 90.29}
}
@misc{darwin180b_rsi_2026,
title = {Darwin-180B-RSI: Recursive Self-Improvement on Verified Answers for a 180B Mixture-of-Experts Reasoning Model},
author = {FINAL-Bench / Darwin Research Team},
year = {2026},
howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI}},
note = {GPQA Diamond 94.44 \% · MMLU-Pro 88.12 \%}
}
@misc{darwin_family_2026,
title = {Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning},
author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Shin, Junghoon},
year = {2026},
eprint = {2605.14386},
archivePrefix = {arXiv}
}
📜 License
Darwin-180B-RSI-R3 is a derivative of Qwen3.8-Flash-Next (through Darwin-180B-RSI) and is distributed under the Qwen Community License 1.0 (see LICENSE).
🏢 About
Built by VIDRAFT · evaluated with FINAL-Bench. Part of the Darwin Family.
