The Darwin family's first
DUO
model — two domain-verified SOTA models served as a single OpenAI-compatible endpoint.
Darwin-60B-DUO unifies two specialist models from the Darwin family behind a single API:
Darwin-28B-REASON
— Hugging Face leaderboard
GPQA Diamond rank #3
, English graduate-level reasoning specialist.
AWAXIS-Think-31B
— National
K-AI Leaderboard rank #1
(operated by the Ministry of Science and ICT of the Republic of Korea), Korean specialist.
A Hybrid-A router automatically dispatches each request to the optimal strategy (single route / sequential collaboration / ensemble), so callers see one model and one endpoint while internally benefiting from both specialists.
Model Description
Darwin-60B-DUO is a
gateway-orchestrated aggregate
of two constituent base models. The repository contains a FastAPI orchestrator, configuration, and Docker Compose recipe. The model weights themselves live in the constituent repositories and are loaded at runtime by two vLLM backends.
Note on AWAXIS membership.
AWAXIS-Think-31B is also part of the Darwin family — it is the Korean specialist branch distilled by the Darwin team on top of Google's Gemma-4 base, complementing the original Qwen3.5-line Darwin lineage as the family's second axis.
Hybrid-A Orchestration
The gateway analyzes each incoming request and selects one of five strategies. The default distribution observed on representative traffic is:
Strategy
When it fires
Backends called
Cost vs. single 30 B
Share
route_awaxis
Korean-dominant input
AWAXIS only
1×
~50 %
route_darwin
English-dominant input
Darwin only
1×
~20 %
split_refine
Korean output requiring rigorous English / STEM reasoning
Darwin (draft) → AWAXIS (polish)
2×
~15 %
split_refine_reverse
English output requiring Korean cultural / linguistic context
AWAXIS (draft) → Darwin (polish)
2×
~5 %
ensemble_v1
Short-answer / multiple-choice queries
Both backends with self-consistency + cross-verification tournament
2×
~10 %
Average effective cost is approximately
1.3× a single 30 B model
— 70 % of traffic is served by a single backend; the remaining 30 % uses both.
Intended Use
Primary use cases
Bilingual Korean-English assistants
that require both Korean fluency and high-quality English reasoning.
Single-endpoint integration
where downstream tooling already targets the OpenAI Chat Completions API (LangChain, LlamaIndex, OpenAI SDK, Continue, Cursor, etc.).
Cost-conscious deployment
— most traffic is served by a single backend at 1× cost while difficult cross-domain queries automatically receive a 2× collaboration.
Out-of-scope
Vision / video generation.
Both constituent models are text-mode only as deployed here (
--limit-mm-per-prompt {"image":0,"video":0}
).
Real-time streaming.
Initial gateway release does not stream token-by-token. Streaming is planned for v1.1.
Direct
AutoModel.from_pretrained()
loading.
This repository contains an orchestrator, not unified weights. Use the gateway or Docker Compose.
git clone https://huggingface.co/FINAL-Bench/Darwin-60B-DUO
cd Darwin-60B-DUO
# HF token is needed only to download constituent weights on first launchexport HF_TOKEN=hf_xxx
docker compose -f docker/docker-compose.yml up -d
# Verify
curl http://localhost:8000/v1/models
# {"object":"list","data":[{"id":"darwin-60b-duo",...}]}
Single-GPU collocation.
With FP8 quantization the combined footprint is ~30 GB. Set both backends to
CUDA_VISIBLE_DEVICES=0
and
--gpu-memory-utilization 0.45
to colocate on a single 80 GB B200 / H100.
OpenAI-compatible call
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="anything")
# The router picks split_refine (Korean output with English/STEM reasoning)
resp = client.chat.completions.create(
model="darwin-60b-duo",
messages=[{
"role": "user",
"content": "Explain the practical difference between GPT-5 and Claude's reasoning, in Korean.",
}],
)
print(resp.choices[0].message.content)
# Darwin produces the English reasoning; AWAXIS polishes it into natural Korean.
You can also force a specific strategy via the non-standard
duo_strategy
field:
resp = client.chat.completions.create(
model="darwin-60b-duo",
messages=[{"role":"user","content":"Which is correct? (A) ... (B) ..."}],
extra_body={"duo_strategy": "ensemble_v1"}, # force MAJ@8 + cross-verify
)
Inspect the chosen strategy in the response under
_duo_route
:
Darwin-60B-DUO sits at the confluence of two complete lineages — the
Qwen3.5-based Darwin lineage
(English reasoning) and the
Gemma-4-based Darwin Korean-specialist branch (AWAXIS)
. The full family tree, with both constituent ancestries fully expanded:
graph TD
%% Lineage A — English reasoning
A1[Cohere Command A+ - 218B foundation]:::found --> A2[Darwin-28B-Opus - English reasoning base]:::parent
A2 --> A3[Darwin-28B-REASON - HF GPQA Diamond #3]:::spec
%% Other Darwin parents
A1 -.-> P1[Darwin-218B-Delphi - cascade flagship GPQA 90.91%]:::parent
A2 -.-> P2[Darwin-9B - omni-modal ko/en compact]:::parent
P2 -.-> P3[Darwin-31B-Opus - Korean multimodal base]:::parent
%% Lineage B — Korean specialist
B1[Google Gemma-4-31B-it - Korean/multilingual base]:::found --> B2[TeichAI gemma-4-31B-it-Claude-Opus-Distill-v2]:::parent
B2 --> B3[AWAXIS-Think-31B - National K-AI Leaderboard #1, Darwin family Korean specialist]:::spec
%% The DUO unification
A3 --> DUO[Darwin-60B-DUO - this model]:::duo
B3 --> DUO
classDef found fill:#e8f0fe,stroke:#1a73e8,color:#0a0a0a
classDef parent fill:#fff4e5,stroke:#f29900,color:#0a0a0a
classDef spec fill:#e6f4ea,stroke:#34a853,color:#0a0a0a
classDef duo fill:#fce8f3,stroke:#d81b60,color:#0a0a0a,stroke-width:3px
Plain-text fallback
Darwin Family
Lineage A — English reasoning (Qwen3.5-line)
└── Cohere Command A+ (218B foundation)
└── Darwin-28B-Opus (English reasoning base)
└── Darwin-28B-REASON 🥉 ← HF GPQA Diamond #3
(English reasoning specialist)
│
│ Related Darwin parents in this lineage:
│ • Darwin-218B-Delphi (cascade flagship, GPQA Diamond 90.91 %)
│ • Darwin-9B (omni-modal ko/en compact)
│ • Darwin-31B-Opus (Korean multimodal base)
Lineage B — Korean specialist (Gemma-4-line)
└── Google Gemma-4-31B-it (Korean / multilingual base)
└── TeichAI gemma-4-31B-it-Claude-Opus-Distill-v2
└── AWAXIS-Think-31B 🥇 ← National K-AI Leaderboard #1
(Darwin family Korean specialist)
│ │
└──────── DUO unification ────┘
↓
⭐ Darwin-60B-DUO ⭐ ← THIS MODEL
"Two SOTAs, one OpenAI-compatible endpoint."
The HF
Model tree
widget (right sidebar) automatically renders the upstream chain from each
base_model
declared in the YAML frontmatter, so the full ancestry — Cohere Command A+ and Google Gemma-4-31B-it at the roots — is browsable directly on this page.
Operation Modes
Mode 1 — Route (single backend, ~70 % of traffic)
The router selects one backend based on language ratio and lightweight keyword heuristics:
korean_ratio(prompt) > 0.3
→ AWAXIS
ASCII / code / math markers (
def
,
import
,
\boxed
,
prove
, …) → Darwin
One model drafts, the other polishes. The polish instruction is language-adaptive:
User: "Explain entropy intuitively in Korean."
Step 1 — Darwin (rigorous English reasoning)
"Entropy quantifies the number of microstates compatible
with a given macrostate, representing disorder ..."
Step 2 — AWAXIS (natural Korean polish)
"엔트로피는 쉽게 말하면 '무질서함의 정도'입니다.
같은 모습으로 보이지만 사실 그 안에 ..."
The reverse path (
AWAXIS draft → Darwin polish
) fires when the output language is English but the prompt requires Korean cultural or linguistic context.
Mode 3 — Ensemble V₁ Tournament (~10 % of traffic)
For multiple-choice and short-answer queries, both backends produce
N = 8 samples
at temperature 0.7. Each backend's answer is its own majority vote (self-consistency). If the two majorities agree → return. If they disagree → each backend verifies the pair (cross-verification), and the tournament winner is selected. A confidence tiebreaker (own-vote count) resolves split verdicts.
National K-AI Leaderboard (Republic of Korea, MSIT)
#1
Darwin-60B-DUO aggregate
Benchmark
Status
GPQA Diamond (full 198 questions)
TBA
KMMLU
TBA
CLIcK (Korean cultural reasoning)
TBA
Helmet / Ruler (long context)
TBA
Needle-in-Haystack 32 K / 128 K
NIAH 32 K: 5/5 per backend (sanity, single model only) — full DUO numbers TBA
Aggregate DUO benchmark results will be published in
benchmarks/
after formal evaluation. The verified constituent ranks above are independent third-party measurements and are not aggregate DUO scores.
Cost / latency notes
Route mode:
comparable to a single 30 B FP8 backend (1× latency, 1× compute).
Ensemble V₁:
~2× compute (parallel) plus a short cross-verify round when majorities disagree.
Bias, Risks, and Limitations
Hallucination.
Standard LLM caveats apply. Both backends can produce confident but incorrect outputs, especially on out-of-distribution queries.
Disagreement bias.
Empirically, the V₁ tournament occasionally selects a wrong answer that both backends collectively favor over a single backend's correct one. The gateway exposes the routing decision in
_duo_route
for auditability.
Language coverage.
Best performance on English and Korean. Other languages fall back to the closer-fit backend without explicit optimization.
Combined weights are not bundled.
The aggregator pulls each backend's weights from the constituent repositories. Network and disk for both is required.
Two-GPU baseline.
BF16 deployment requires two GPUs. FP8 quantization enables single-GPU operation on B200 / H100 (80 GB).
Training data cut-off.
Darwin-28B-REASON: ~ 2026-Q1. AWAXIS-Think-31B: ~ 2026-Q1.
License
Darwin-60B-DUO inherits the
Gemma Terms of Use
as its effective combined license — the more restrictive of the two constituent licenses.
Constituent
License
Darwin-28B-REASON
Apache-2.0
AWAXIS-Think-31B
Gemma (inherited from Gemma-4)
Darwin-60B-DUO (aggregate)
Gemma
(combined-license inheritance)
The orchestrator code (
gateway/
,
docker/
) is offered under Apache-2.0 to maximize developer flexibility; combined-license inheritance applies to served model behavior only.
Issues and discussions: please open a thread on the
Community
tab of this repository.
Citation
@misc{darwin60b-duo-2026,
title = {Darwin-60B-DUO: A single-endpoint DUO of an English-reasoning SOTA
and a Korean SOTA via Hybrid-A orchestration},
author = {FINAL-Bench Team and Anserwise Team and VIDRAFT},
year = {2026},
howpublished = {Hugging Face},
url = {https://huggingface.co/FINAL-Bench/Darwin-60B-DUO}
}
Runs of FINAL-Bench Darwin-60B-DUO on huggingface.co
118
Total runs
1
24-hour runs
15
3-day runs
35
7-day runs
86
30-day runs
More Information About Darwin-60B-DUO huggingface.co Model
Darwin-60B-DUO huggingface.co is an AI model on huggingface.co that provides Darwin-60B-DUO's model effect (), which can be used instantly with this FINAL-Bench Darwin-60B-DUO model. huggingface.co supports a free trial of the Darwin-60B-DUO model, and also provides paid use of the Darwin-60B-DUO. Support call Darwin-60B-DUO model through api, including Node.js, Python, http.
Darwin-60B-DUO huggingface.co is an online trial and call api platform, which integrates Darwin-60B-DUO's modeling effects, including api services, and provides a free online trial of Darwin-60B-DUO, you can try Darwin-60B-DUO online for free by clicking the link below.
FINAL-Bench Darwin-60B-DUO online free url in huggingface.co:
Darwin-60B-DUO is an open source model from GitHub that offers a free installation service, and any user can find Darwin-60B-DUO on GitHub to install. At the same time, huggingface.co provides the effect of Darwin-60B-DUO install, users can directly use Darwin-60B-DUO installed effect in huggingface.co for debugging and trial. It also supports api for free installation.