FINAL-Bench / Darwin-180B-RSI-R3

huggingface.co
Total runs: 58
24-hour runs: 55
7-day runs: 58
30-day runs: 58
Model's Last Updated: October 02 2026
image-text-to-text

Introduction of Darwin-180B-RSI-R3

Model Details of Darwin-180B-RSI-R3

Darwin-180B-RSI-R3

The second round of model-level self-improvement on top of Darwin-180B-RSI .

Darwin-180B-RSI (R1) holds first place on seven Hugging Face official leaderboards. R3 continues training from R1 with the same recipe: the model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used.

R3 is released so anyone can download it, run it, and check the numbers below.

What changed from R1
  • Starting point: R1 weights (not the parent). R3 is a true second round.
  • Practice problems: 3,000 SuperGPQA questions (middle and hard difficulty) that were never used in R1 training. R1 solved each one 8 times.
  • What it learned from: only the "boundary" problems, where R1 was right on 2 to 6 of 8 attempts (462 problems). From those, up to 2 of R1's own correct, untruncated solutions per problem (714 solutions in total).
  • What was trained: the same components as R1 (attention paths and shared experts) via LoRA, then merged. All 512 routed experts, the router and the vision encoder are unchanged.
  • No benchmark data: GPQA Diamond and the held-out set below were never used for training or selection.
Results (same settings for every model)
Held-out SuperGPQA, 1,000 questions never used in training or selection

4 samples per question, 16K thinking budget, temperature 1.0.

Model Single sample Mean of 4 Majority of 4
R1 (Darwin-180B-RSI) 65.30 65.67 68.30
R3 (this model) 66.30 66.70 69.00

Paired per-question difference, R1 → R3 (mean of 4): +1.03 points, 95% CI [+0.05, +2.00] . The gain is small but statistically significant.

GPQA Diamond, 198 questions

8 samples per question, 32K thinking budget, temperature 1.0.

Model Single sample Mean of 8 Majority of 8
R0 (parent, Qwen3.8-Flash-Next) 84.85 85.35 90.91
R1 (Darwin-180B-RSI) 84.85 85.80 89.90
R3 (this model) 85.86 86.05 90.40

Paired differences on GPQA (mean of 8): R0 → R3 +0.69 [−0.71, +2.10], R1 → R3 +0.25 [−1.18, +1.69]. With 198 questions these are within noise; we report them as measured.

How to read this: on a large held-out set, the second round of self-improvement produced a measurable gain over R1. On GPQA the model was already near its ceiling and the differences are not significant.

Leaderboard status

The seven Hugging Face leaderboard #1 results belong to R1 ( Darwin-180B-RSI ). R3 has not been submitted to any leaderboard. We are asking independent evaluators to measure it directly.

Quickstart

Serving is the same as Darwin-180B-RSI (vLLM, tensor parallel + expert parallel). This is a reasoning model; give it a long generation budget.

vllm serve FINAL-Bench/Darwin-180B-RSI-R3 \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 139264

Recommended sampling: temperature 1.0, top_p 0.95, top_k 20, up to 131,072 generated tokens.

ZTC

The ZTC probe published with Darwin-180B-RSI was fitted on R1's hidden states. A probe fitted on R3 will be added to this repository; until then, use R1's probe only as a rough signal.

Model-level RSI vs. harness-level RSI

Darwin-180B-RSI-R3 is Model-level RSI : the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. Harness-level RSI (e.g., Google's RRSI) improves the prompts, tools and workflow around a fixed model. The two are complementary.

License

Qwen Community License 1.0 (inherited from Qwen3.8-Flash-Next). See LICENSE .

About

Built by VIDRAFT . Darwin family paper: arXiv 2605.14386 .

Runs of FINAL-Bench Darwin-180B-RSI-R3 on huggingface.co

58
Total runs
55
24-hour runs
58
3-day runs
58
7-day runs
58
30-day runs

More Information About Darwin-180B-RSI-R3 huggingface.co Model

More Darwin-180B-RSI-R3 license Visit here:

https://choosealicense.com/licenses/qwen-community-1.0

Darwin-180B-RSI-R3 huggingface.co

Darwin-180B-RSI-R3 huggingface.co is an AI model on huggingface.co that provides Darwin-180B-RSI-R3's model effect (), which can be used instantly with this FINAL-Bench Darwin-180B-RSI-R3 model. huggingface.co supports a free trial of the Darwin-180B-RSI-R3 model, and also provides paid use of the Darwin-180B-RSI-R3. Support call Darwin-180B-RSI-R3 model through api, including Node.js, Python, http.

Darwin-180B-RSI-R3 huggingface.co Url

https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3

FINAL-Bench Darwin-180B-RSI-R3 online free

Darwin-180B-RSI-R3 huggingface.co is an online trial and call api platform, which integrates Darwin-180B-RSI-R3's modeling effects, including api services, and provides a free online trial of Darwin-180B-RSI-R3, you can try Darwin-180B-RSI-R3 online for free by clicking the link below.

FINAL-Bench Darwin-180B-RSI-R3 online free url in huggingface.co:

https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3

Darwin-180B-RSI-R3 install

Darwin-180B-RSI-R3 is an open source model from GitHub that offers a free installation service, and any user can find Darwin-180B-RSI-R3 on GitHub to install. At the same time, huggingface.co provides the effect of Darwin-180B-RSI-R3 install, users can directly use Darwin-180B-RSI-R3 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Darwin-180B-RSI-R3 install url in huggingface.co:

https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3

Url of Darwin-180B-RSI-R3

Darwin-180B-RSI-R3 huggingface.co Url

Provider of Darwin-180B-RSI-R3 huggingface.co

FINAL-Bench
ORGANIZATIONS

Other API from FINAL-Bench