180B Mixture-of-Experts · vision-language ·
#1 on five Hugging Face official leaderboards
— AIME 2026 100 · HMMT Feb 2026 100 · GPQA Diamond 94.44 · MMLU-Pro 88.12 · MMMU-Pro 79.48 ·
self-improving
reasoning
·
MoE 512 experts
·
262K long context
·
image + text
·
Korean + English
·
self-improvement
·
ZTC
The newest flagship of the Darwin family — #1 on AIME 2026, HMMT Feb 2026, GPQA Diamond, MMLU-Pro and MMMU-Pro,
and a model that gets better by learning from its own verified work.
🏆 Five #1s — head-to-head with Chinese frontier models
Scores as listed on the Hugging Face
official
benchmark leaderboards (self-reported by each model's publisher).
🥇 = #1 on that leaderboard. "—" = not reported.
Model
AIME 2026
GPQA Diamond
MMLU-Pro
MMMU-Pro
HMMT Feb 2026
🧬 Darwin-180B-RSI (ours · 🇰🇷)
100
🥇
94.44
🥇
88.12
🥇
79.48
🥇
100
🥇
Kimi-K3 (Moonshot AI)
—
93.5
—
—
—
Kimi-K2.6 (Moonshot AI)
96.4
90.5
—
79.4
92.7
DeepSeek-V4-Pro (DeepSeek)
—
90.1
87.5
—
—
Qwen3.5-397B-A17B (Alibaba)
93.33
88.4
87.8
—
87.88
MiniMax-M2.1 (MiniMax)
—
80.81
88
—
—
GLM-5 (Zhipu AI)
95.83
86
86
—
86.36
Intern-S2-Preview (Shanghai AI Lab)
—
—
88
76.88
87.31
Step-3.5-Flash (StepFun)
96.67
83.5
84.4
—
86.36
Leaderboard values are each publisher's own reported numbers; settings (samples, voting, thinking budget) differ across models. Darwin-180B-RSI settings are listed in the evaluation protocol below.
🧬 The Darwin Family
Darwin
is
VIDRAFT
's measurement-driven reasoning model family —
50+ official models
,
400+ community derivatives
, and now
two places in the GPQA Diamond top 3
(Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).
🧬 Darwin — evolve the parent, keep what works
Darwin treats a strong open model as a
parent
. It measures where the parent is weak,
and strengthens exactly those parts — instead of re-training everything and risking what already works.
Diagnose before you change.
Every Darwin generation starts from a measured weakness map of the parent.
Change little, precisely.
Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
Proven capability over new guesses.
Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient —
the model's own verified work
.
Measured, not claimed.
Every change must beat the parent on held-out tests before it ships.
Model
Scale
GPQA Diamond
Darwin-9B-NEG
9B
84.3
Darwin-27B-Opus
27B dense
86.9
Darwin-36B-Opus
36B MoE
88.4
Darwin-28B-REASON
28B + DELPHI
89.39
Darwin-397B-ZTC
397B MoE (FP8)
93.43
Darwin-180B-RSI
180B MoE
94.44
Lineage
Role
Parent
Qwen/Qwen3.8-Flash-Next
180B MoE vision-language backbone · Qwen Community License 1.0
Darwin RSI
self-improvement on verified answers
the parent's own solutions, checked against verifiable answer keys, fed back as training signal
Preserved
512 routed experts · router · vision encoder
untouched — the parent's knowledge stays intact
ZTC
zero-token confidence readout
see below
📄 Darwin Platform & Research
Darwin Family
— MRI trust-weighted evolutionary merging for training-free scaling of language-model reasoning (
arXiv:2605.14386
)
Placement Is Free, Composition Is Not
— the Latin square as a provably-balanced construction for heterogeneous sequence-mixer stacks (
2609.20269
) — the AETHER architecture line
FINAL Bench
— VIDRAFT's measurement-driven evaluation framework (SSRN)
Recursive self-improvement (RSI)
is the core of this generation.
Instead of distilling a bigger teacher, the model improves by learning from itself:
Solve
— the model works through practice problems it has never seen in evaluation.
Verify
— its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
Learn
— it is re-trained on the reasoning that turned out to be correct.
Repeat
— the improved model becomes the next solver.
What it bought in this release:
Parent (Qwen3.8-Flash-Next)
Darwin-180B-RSI
Average reasoning length (MMLU-Pro)
4,320 tokens
3,833 tokens (−11 %)
MMLU-Pro accuracy
88.04 %
88.12 %
Same or better accuracy with shorter reasoning
— cheaper and faster to serve.
Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).
🏛️ ZTC — it knows before it answers
Zero-Token Confidence (ZTC)
reads the model's own internal state
once, before generation
,
and returns the probability that the answer it is about to give is correct —
no extra tokens, no second model.
All numbers are self-measured and reproducible with the settings above. Majority-vote scores are system scores (several samples per question) and are labeled as such.
MMLU-Pro by category (single sample)
— strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.
from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)
Transformers
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
Tip:
this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning.
Short budgets truncate the reasoning and cost accuracy.
⚠️ Limitations and disclosure
Scores are self-measured with the settings stated in the Results table; majority-vote numbers use several samples per question.
Very long reasoning is normal for hard problems; a short thinking budget will truncate answers and lower accuracy.
Like every LLM, the model can be confidently wrong — use the ZTC confidence readout to gate high-stakes actions.
@misc{darwin180b_rsi_2026,
title = {Darwin-180B-RSI: Recursive Self-Improvement on Verified Answers for a 180B Mixture-of-Experts Reasoning Model},
author = {FINAL-Bench / Darwin Research Team},
year = {2026},
howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI}},
note = {GPQA Diamond 94.44 \% · MMLU-Pro 88.12 \%}
}
@misc{darwin_family_2026,
title = {Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning},
author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Shin, Junghoon},
year = {2026},
eprint = {2605.14386},
archivePrefix = {arXiv}
}
@misc{latin_square_2026,
title = {Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks},
author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Kim, Minseo},
year = {2026},
eprint = {2609.20269},
archivePrefix = {arXiv}
}
📜 License
Darwin-180B-RSI is a derivative of
Qwen3.8-Flash-Next
and is distributed under the
Qwen Community License 1.0
(see
LICENSE
).
Darwin-180B-RSI huggingface.co is an AI model on huggingface.co that provides Darwin-180B-RSI's model effect (), which can be used instantly with this FINAL-Bench Darwin-180B-RSI model. huggingface.co supports a free trial of the Darwin-180B-RSI model, and also provides paid use of the Darwin-180B-RSI. Support call Darwin-180B-RSI model through api, including Node.js, Python, http.
Darwin-180B-RSI huggingface.co is an online trial and call api platform, which integrates Darwin-180B-RSI's modeling effects, including api services, and provides a free online trial of Darwin-180B-RSI, you can try Darwin-180B-RSI online for free by clicking the link below.
FINAL-Bench Darwin-180B-RSI online free url in huggingface.co:
Darwin-180B-RSI is an open source model from GitHub that offers a free installation service, and any user can find Darwin-180B-RSI on GitHub to install. At the same time, huggingface.co provides the effect of Darwin-180B-RSI install, users can directly use Darwin-180B-RSI installed effect in huggingface.co for debugging and trial. It also supports api for free installation.