FINAL-Bench / Darwin-180B-RSI

huggingface.co
Total runs: 546
24-hour runs: 62
7-day runs: 476
30-day runs: 476
Model's Last Updated: October 02 2026
image-text-to-text

Introduction of Darwin-180B-RSI

Model Details of Darwin-180B-RSI

Darwin-180B-RSI

180B Mixture-of-Experts · vision-language · #1 on five Hugging Face official leaderboards — AIME 2026 100 · HMMT Feb 2026 100 · GPQA Diamond 94.44 · MMLU-Pro 88.12 · MMMU-Pro 79.48 · self-improving

reasoning · MoE 512 experts · 262K long context · image + text · Korean + English · self-improvement · ZTC

The newest flagship of the Darwin family — #1 on AIME 2026, HMMT Feb 2026, GPQA Diamond, MMLU-Pro and MMMU-Pro, and a model that gets better by learning from its own verified work.


🏆 Five #1s — head-to-head with Chinese frontier models

Darwin-180B-RSI vs Chinese frontier models

Five leaderboards — full field

Scores as listed on the Hugging Face official benchmark leaderboards (self-reported by each model's publisher). 🥇 = #1 on that leaderboard. "—" = not reported.

Model AIME 2026 GPQA Diamond MMLU-Pro MMMU-Pro HMMT Feb 2026
🧬 Darwin-180B-RSI (ours · 🇰🇷) 100 🥇 94.44 🥇 88.12 🥇 79.48 🥇 100 🥇
Kimi-K3 (Moonshot AI) — 93.5 — — —
Kimi-K2.6 (Moonshot AI) 96.4 90.5 — 79.4 92.7
DeepSeek-V4-Pro (DeepSeek) — 90.1 87.5 — —
Qwen3.5-397B-A17B (Alibaba) 93.33 88.4 87.8 — 87.88
MiniMax-M2.1 (MiniMax) — 80.81 88 — —
GLM-5 (Zhipu AI) 95.83 86 86 — 86.36
Intern-S2-Preview (Shanghai AI Lab) — — 88 76.88 87.31
Step-3.5-Flash (StepFun) 96.67 83.5 84.4 — 86.36

Leaderboard values are each publisher's own reported numbers; settings (samples, voting, thinking budget) differ across models. Darwin-180B-RSI settings are listed in the evaluation protocol below.


🧬 The Darwin Family

Darwin is VIDRAFT 's measurement-driven reasoning model family — 50+ official models , 400+ community derivatives , and now two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).


🧬 Darwin — evolve the parent, keep what works

Darwin treats a strong open model as a parent . It measures where the parent is weak, and strengthens exactly those parts — instead of re-training everything and risking what already works.

  • Diagnose before you change. Every Darwin generation starts from a measured weakness map of the parent.
  • Change little, precisely. Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
  • Proven capability over new guesses. Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — the model's own verified work .
  • Measured, not claimed. Every change must beat the parent on held-out tests before it ships.
Model Scale GPQA Diamond
Darwin-9B-NEG 9B 84.3
Darwin-27B-Opus 27B dense 86.9
Darwin-36B-Opus 36B MoE 88.4
Darwin-28B-REASON 28B + DELPHI 89.39
Darwin-397B-ZTC 397B MoE (FP8) 93.43
Darwin-180B-RSI 180B MoE 94.44
Lineage
Role
Parent Qwen/Qwen3.8-Flash-Next 180B MoE vision-language backbone · Qwen Community License 1.0
Darwin RSI self-improvement on verified answers the parent's own solutions, checked against verifiable answer keys, fed back as training signal
Preserved 512 routed experts · router · vision encoder untouched — the parent's knowledge stays intact
ZTC zero-token confidence readout see below

📄 Darwin Platform & Research
  • Darwin Family — MRI trust-weighted evolutionary merging for training-free scaling of language-model reasoning ( arXiv:2605.14386 )
  • Placement Is Free, Composition Is Not — the Latin square as a provably-balanced construction for heterogeneous sequence-mixer stacks ( 2609.20269 ) — the AETHER architecture line
  • FINAL Bench — VIDRAFT's measurement-driven evaluation framework (SSRN)
  • Four-layer Pre-AGI roadmap — Darwin → AETHER → PROMETHEUS → HEPHAESTUS
  • Collections: Darwin Family · ZTC Models — JEV ecosystems

🔁 RSI — a model that improves from its own work

Recursive self-improvement (RSI) is the core of this generation. Instead of distilling a bigger teacher, the model improves by learning from itself:

  1. Solve — the model works through practice problems it has never seen in evaluation.
  2. Verify — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
  3. Learn — it is re-trained on the reasoning that turned out to be correct.
  4. Repeat — the improved model becomes the next solver.

What it bought in this release:

Parent (Qwen3.8-Flash-Next) Darwin-180B-RSI
Average reasoning length (MMLU-Pro) 4,320 tokens 3,833 tokens (−11 %)
MMLU-Pro accuracy 88.04 % 88.12 %

Same or better accuracy with shorter reasoning — cheaper and faster to serve. Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).


🏛️ ZTC — it knows before it answers

Zero-Token Confidence (ZTC) reads the model's own internal state once, before generation , and returns the probability that the answer it is about to give is correct — no extra tokens, no second model.

{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}

Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know".

What ships with this model

File Role
handler.py one call returns the answer and its confidence as JSON
ztc/ztc_probe_darwin180rsi.npz the ZTC readout for this model (final layer, last prompt token)
ztc/usage.py minimal example
from handler import EndpointHandler
h = EndpointHandler("FINAL-Bench/Darwin-180B-RSI")   # local snapshot path
print(h({"inputs": "What is 17 * 23?"}))
# [{"answer": "...391...", "confidence": 0.97, "ztc_score": 2.1, "truncated": false}]

Same format as Darwin-397B-ZTC . The readout is fitted only on practice data that is disjoint from every benchmark reported here.


🏆 Results
Benchmark Score Setting Leaderboard
GPQA Diamond (198) 94.44 majority vote over up to 16 samples · 131,072-token thinking budget #1
MMLU-Pro (12,032) 88.12 single sample · 131,072-token thinking budget #1
AIME 2026 (30) 100.0 majority vote over 16 samples (mean accuracy 98.75) · 131,072-token thinking budget #1
HMMT Feb 2026 (33) 100.0 majority vote over 16 samples (mean accuracy 96.59) · 131,072-token thinking budget #1
MMMU-Pro (vision, 1,730) 79.48 majority vote over 3 samples · 131,072-token thinking budget #1
Evaluation protocol

Common to every benchmark

Setting Value
Thinking budget 131,072 tokens (max generated tokens per sample)
Sampling temperature 1.0 · top_p 0.95 · top_k 20
Precision bf16
Engine vLLM, tensor parallel 8 (or 4), expert parallel

Per benchmark

Benchmark Samples per question Reported score
AIME 2026 16 majority vote (maj@16); mean over 16 = 98.75
HMMT Feb 2026 16 majority vote (maj@16); mean over 16 = 96.59
GPQA Diamond up to 16 majority vote
MMLU-Pro 1 single sample (no voting)
MMMU-Pro (vision) 3 majority vote (maj@3)

All numbers are self-measured and reproducible with the settings above. Majority-vote scores are system scores (several samples per question) and are labeled as such.

MMLU-Pro by category (single sample) — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.


⚙️ Specifications
Architecture Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers)
Layers / hidden 48 / 2,560
Experts 512 routed (10 active per token) + shared expert
Context 262,144 tokens
Vocabulary 248,320
Modalities image + text → text
Precision bf16 (~336 GB)

🚀 Quickstart
Serving with vLLM (8 × B200 or equivalent)
vllm serve FINAL-Bench/Darwin-180B-RSI \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 135168 --trust-remote-code
Chat Completions (OpenAI-compatible)
from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
    messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
    temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)
Transformers
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

Tip: this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning. Short budgets truncate the reasoning and cost accuracy.


⚠️ Limitations and disclosure
  • Scores are self-measured with the settings stated in the Results table; majority-vote numbers use several samples per question.
  • Very long reasoning is normal for hard problems; a short thinking budget will truncate answers and lower accuracy.
  • Like every LLM, the model can be confidently wrong — use the ZTC confidence readout to gate high-stakes actions.

🔗 Related Darwin Models

📚 Citation
@misc{darwin180b_rsi_2026,
  title  = {Darwin-180B-RSI: Recursive Self-Improvement on Verified Answers for a 180B Mixture-of-Experts Reasoning Model},
  author = {FINAL-Bench / Darwin Research Team},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI}},
  note   = {GPQA Diamond 94.44 \% · MMLU-Pro 88.12 \%}
}

@misc{darwin_family_2026,
  title  = {Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Shin, Junghoon},
  year   = {2026},
  eprint = {2605.14386},
  archivePrefix = {arXiv}
}

@misc{latin_square_2026,
  title  = {Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Kim, Minseo},
  year   = {2026},
  eprint = {2609.20269},
  archivePrefix = {arXiv}
}

📜 License

Darwin-180B-RSI is a derivative of Qwen3.8-Flash-Next and is distributed under the Qwen Community License 1.0 (see LICENSE ).

🏢 About

Built by VIDRAFT · evaluated with FINAL-Bench .

This model is part of the Darwin Family .

Runs of FINAL-Bench Darwin-180B-RSI on huggingface.co

546
Total runs
62
24-hour runs
250
3-day runs
476
7-day runs
476
30-day runs

More Information About Darwin-180B-RSI huggingface.co Model

More Darwin-180B-RSI license Visit here:

https://choosealicense.com/licenses/qwen-community-1.0

Darwin-180B-RSI huggingface.co

Darwin-180B-RSI huggingface.co is an AI model on huggingface.co that provides Darwin-180B-RSI's model effect (), which can be used instantly with this FINAL-Bench Darwin-180B-RSI model. huggingface.co supports a free trial of the Darwin-180B-RSI model, and also provides paid use of the Darwin-180B-RSI. Support call Darwin-180B-RSI model through api, including Node.js, Python, http.

FINAL-Bench Darwin-180B-RSI online free

Darwin-180B-RSI huggingface.co is an online trial and call api platform, which integrates Darwin-180B-RSI's modeling effects, including api services, and provides a free online trial of Darwin-180B-RSI, you can try Darwin-180B-RSI online for free by clicking the link below.

FINAL-Bench Darwin-180B-RSI online free url in huggingface.co:

https://huggingface.co/FINAL-Bench/Darwin-180B-RSI

Darwin-180B-RSI install

Darwin-180B-RSI is an open source model from GitHub that offers a free installation service, and any user can find Darwin-180B-RSI on GitHub to install. At the same time, huggingface.co provides the effect of Darwin-180B-RSI install, users can directly use Darwin-180B-RSI installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Darwin-180B-RSI install url in huggingface.co:

https://huggingface.co/FINAL-Bench/Darwin-180B-RSI

Url of Darwin-180B-RSI

Darwin-180B-RSI huggingface.co Url

Provider of Darwin-180B-RSI huggingface.co

FINAL-Bench
ORGANIZATIONS

Other API from FINAL-Bench