FINAL-Bench / Darwin-60B-DUO

huggingface.co
Total runs: 118
24-hour runs: 1
7-day runs: 35
30-day runs: 86
Model's Last Updated: July 23 2026
text-generation

Introduction of Darwin-60B-DUO

Model Details of Darwin-60B-DUO

Darwin-60B-DUO

The Darwin family's first DUO model — two domain-verified SOTA models served as a single OpenAI-compatible endpoint.

Darwin-60B-DUO unifies two specialist models from the Darwin family behind a single API:

  • Darwin-28B-REASON — Hugging Face leaderboard GPQA Diamond rank #3 , English graduate-level reasoning specialist.
  • AWAXIS-Think-31B — National K-AI Leaderboard rank #1 (operated by the Ministry of Science and ICT of the Republic of Korea), Korean specialist.

A Hybrid-A router automatically dispatches each request to the optimal strategy (single route / sequential collaboration / ensemble), so callers see one model and one endpoint while internally benefiting from both specialists.


Model Description

Darwin-60B-DUO is a gateway-orchestrated aggregate of two constituent base models. The repository contains a FastAPI orchestrator, configuration, and Docker Compose recipe. The model weights themselves live in the constituent repositories and are loaded at runtime by two vLLM backends.

Component Source Architecture Parameters Verified Rank
English reasoning specialist FINAL-Bench/Darwin-28B-REASON Qwen3.5 multimodal 26.9 B HF GPQA Diamond #3
Korean specialist (Darwin family, Gemma-4 branch) Anserwise/AWAXIS-Think-31B Gemma-4 multimodal 31.27 B National K-AI Leaderboard #1
Aggregate This repository ( FINAL-Bench/Darwin-60B-DUO ) DUO orchestrator 58.17 B —

Note on AWAXIS membership. AWAXIS-Think-31B is also part of the Darwin family — it is the Korean specialist branch distilled by the Darwin team on top of Google's Gemma-4 base, complementing the original Qwen3.5-line Darwin lineage as the family's second axis.

Hybrid-A Orchestration

The gateway analyzes each incoming request and selects one of five strategies. The default distribution observed on representative traffic is:

Strategy When it fires Backends called Cost vs. single 30 B Share
route_awaxis Korean-dominant input AWAXIS only 1× ~50 %
route_darwin English-dominant input Darwin only 1× ~20 %
split_refine Korean output requiring rigorous English / STEM reasoning Darwin (draft) → AWAXIS (polish) 2× ~15 %
split_refine_reverse English output requiring Korean cultural / linguistic context AWAXIS (draft) → Darwin (polish) 2× ~5 %
ensemble_v1 Short-answer / multiple-choice queries Both backends with self-consistency + cross-verification tournament 2× ~10 %

Average effective cost is approximately 1.3× a single 30 B model — 70 % of traffic is served by a single backend; the remaining 30 % uses both.


Intended Use
Primary use cases
  • Bilingual Korean-English assistants that require both Korean fluency and high-quality English reasoning.
  • Single-endpoint integration where downstream tooling already targets the OpenAI Chat Completions API (LangChain, LlamaIndex, OpenAI SDK, Continue, Cursor, etc.).
  • Cost-conscious deployment — most traffic is served by a single backend at 1× cost while difficult cross-domain queries automatically receive a 2× collaboration.
Out-of-scope
  • Vision / video generation. Both constituent models are text-mode only as deployed here ( --limit-mm-per-prompt {"image":0,"video":0} ).
  • Real-time streaming. Initial gateway release does not stream token-by-token. Streaming is planned for v1.1.
  • Direct AutoModel.from_pretrained() loading. This repository contains an orchestrator, not unified weights. Use the gateway or Docker Compose.
  • Safety-critical decision making (medical diagnosis, legal advice, autonomous control). LLM hallucinations apply.

How to Use
Option A — Docker Compose (recommended)
git clone https://huggingface.co/FINAL-Bench/Darwin-60B-DUO
cd Darwin-60B-DUO

# HF token is needed only to download constituent weights on first launch
export HF_TOKEN=hf_xxx
docker compose -f docker/docker-compose.yml up -d

# Verify
curl http://localhost:8000/v1/models
# {"object":"list","data":[{"id":"darwin-60b-duo",...}]}
Option B — Manual launch (2 × B200 / H100, FP8)
# 1. Darwin-28B-REASON on GPU 0
CUDA_VISIBLE_DEVICES=0 VLLM_DP_MASTER_PORT=45011 \
  vllm serve FINAL-Bench/Darwin-28B-REASON \
    --port 8021 --served-model-name darwin-28r \
    --quantization fp8 --enforce-eager \
    --limit-mm-per-prompt '{"image":0,"video":0}' &

# 2. AWAXIS-Think-31B on GPU 1
CUDA_VISIBLE_DEVICES=1 VLLM_DP_MASTER_PORT=45012 \
  vllm serve Anserwise/AWAXIS-Think-31B \
    --port 8022 --served-model-name awaxis-31b \
    --quantization fp8 --enforce-eager \
    --limit-mm-per-prompt '{"image":0,"video":0}' &

# 3. Gateway
pip install -r gateway/requirements.txt
python gateway/server.py --port 8000 \
    --darwin-url http://127.0.0.1:8021/v1 \
    --awaxis-url http://127.0.0.1:8022/v1

Single-GPU collocation. With FP8 quantization the combined footprint is ~30 GB. Set both backends to CUDA_VISIBLE_DEVICES=0 and --gpu-memory-utilization 0.45 to colocate on a single 80 GB B200 / H100.

OpenAI-compatible call
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="anything")

# The router picks split_refine (Korean output with English/STEM reasoning)
resp = client.chat.completions.create(
    model="darwin-60b-duo",
    messages=[{
        "role": "user",
        "content": "Explain the practical difference between GPT-5 and Claude's reasoning, in Korean.",
    }],
)
print(resp.choices[0].message.content)
# Darwin produces the English reasoning; AWAXIS polishes it into natural Korean.

You can also force a specific strategy via the non-standard duo_strategy field:

resp = client.chat.completions.create(
    model="darwin-60b-duo",
    messages=[{"role":"user","content":"Which is correct? (A) ... (B) ..."}],
    extra_body={"duo_strategy": "ensemble_v1"},  # force MAJ@8 + cross-verify
)

Inspect the chosen strategy in the response under _duo_route :

{
  "choices": [{"message": {"role":"assistant", "content":"..."}}],
  "_duo_route": {
    "strategy": "split_refine",
    "reason": "korean_output_with_english_reasoning",
    "elapsed_s": 4.83,
    "language_ratio": 0.42
  }
}

Darwin Family

Darwin-60B-DUO sits at the confluence of two complete lineages — the Qwen3.5-based Darwin lineage (English reasoning) and the Gemma-4-based Darwin Korean-specialist branch (AWAXIS) . The full family tree, with both constituent ancestries fully expanded:

graph TD
    %% Lineage A — English reasoning
    A1[Cohere Command A+ - 218B foundation]:::found --> A2[Darwin-28B-Opus - English reasoning base]:::parent
    A2 --> A3[Darwin-28B-REASON - HF GPQA Diamond #3]:::spec

    %% Other Darwin parents
    A1 -.-> P1[Darwin-218B-Delphi - cascade flagship GPQA 90.91%]:::parent
    A2 -.-> P2[Darwin-9B - omni-modal ko/en compact]:::parent
    P2 -.-> P3[Darwin-31B-Opus - Korean multimodal base]:::parent

    %% Lineage B — Korean specialist
    B1[Google Gemma-4-31B-it - Korean/multilingual base]:::found --> B2[TeichAI gemma-4-31B-it-Claude-Opus-Distill-v2]:::parent
    B2 --> B3[AWAXIS-Think-31B - National K-AI Leaderboard #1, Darwin family Korean specialist]:::spec

    %% The DUO unification
    A3 --> DUO[Darwin-60B-DUO - this model]:::duo
    B3 --> DUO

    classDef found fill:#e8f0fe,stroke:#1a73e8,color:#0a0a0a
    classDef parent fill:#fff4e5,stroke:#f29900,color:#0a0a0a
    classDef spec fill:#e6f4ea,stroke:#34a853,color:#0a0a0a
    classDef duo fill:#fce8f3,stroke:#d81b60,color:#0a0a0a,stroke-width:3px
Plain-text fallback
Darwin Family

Lineage A — English reasoning (Qwen3.5-line)
└── Cohere Command A+ (218B foundation)
    └── Darwin-28B-Opus (English reasoning base)
        └── Darwin-28B-REASON  🥉  ← HF GPQA Diamond #3
                                     (English reasoning specialist)
    │
    │   Related Darwin parents in this lineage:
    │   • Darwin-218B-Delphi  (cascade flagship, GPQA Diamond 90.91 %)
    │   • Darwin-9B           (omni-modal ko/en compact)
    │   • Darwin-31B-Opus     (Korean multimodal base)

Lineage B — Korean specialist (Gemma-4-line)
└── Google Gemma-4-31B-it (Korean / multilingual base)
    └── TeichAI gemma-4-31B-it-Claude-Opus-Distill-v2
        └── AWAXIS-Think-31B  🥇  ← National K-AI Leaderboard #1
                                     (Darwin family Korean specialist)

         │                             │
         └──────── DUO unification ────┘
                          ↓
              ⭐ Darwin-60B-DUO ⭐  ← THIS MODEL
              "Two SOTAs, one OpenAI-compatible endpoint."

The HF Model tree widget (right sidebar) automatically renders the upstream chain from each base_model declared in the YAML frontmatter, so the full ancestry — Cohere Command A+ and Google Gemma-4-31B-it at the roots — is browsable directly on this page.


Operation Modes
Mode 1 — Route (single backend, ~70 % of traffic)

The router selects one backend based on language ratio and lightweight keyword heuristics:

  • korean_ratio(prompt) > 0.3 → AWAXIS
  • ASCII / code / math markers ( def , import , \boxed , prove , …) → Darwin
  • Mixed → AWAXIS (Korean-first default)
Mode 2 — Split / Refine (sequential collaboration, ~20 % of traffic)

One model drafts, the other polishes. The polish instruction is language-adaptive:

User: "Explain entropy intuitively in Korean."

Step 1 — Darwin (rigorous English reasoning)
        "Entropy quantifies the number of microstates compatible
         with a given macrostate, representing disorder ..."

Step 2 — AWAXIS (natural Korean polish)
        "엔트로피는 쉽게 말하면 '무질서함의 정도'입니다.
         같은 모습으로 보이지만 사실 그 안에 ..."

The reverse path ( AWAXIS draft → Darwin polish ) fires when the output language is English but the prompt requires Korean cultural or linguistic context.

Mode 3 — Ensemble V₁ Tournament (~10 % of traffic)

For multiple-choice and short-answer queries, both backends produce N = 8 samples at temperature 0.7. Each backend's answer is its own majority vote (self-consistency). If the two majorities agree → return. If they disagree → each backend verifies the pair (cross-verification), and the tournament winner is selected. A confidence tiebreaker (own-vote count) resolves split verdicts.


Repository Layout
Darwin-60B-DUO/
├── README.md                  ← this model card
├── config.json                ← DUO configuration & orchestration metadata
├── tokenizer_info.json        ← constituent tokenizer references
├── LICENSE                    ← Gemma + Apache-2.0 dual notice
├── gateway/
│   ├── server.py              ← FastAPI OpenAI-compatible orchestrator
│   ├── router.py              ← language / domain / MCQ classifier
│   ├── refine.py              ← sequential refine (drafter → polisher)
│   ├── ensemble.py            ← V₁ MAJ@N + cross-verification
│   └── requirements.txt
├── docker/
│   └── docker-compose.yml     ← vLLM ×2 + gateway integrated launcher
└── benchmarks/
    └── README.md              ← evaluation roadmap (results TBA)

Evaluation
Verified constituent scores
Constituent Benchmark Rank
Darwin-28B-REASON Hugging Face GPQA Diamond #3
AWAXIS-Think-31B National K-AI Leaderboard (Republic of Korea, MSIT) #1
Darwin-60B-DUO aggregate
Benchmark Status
GPQA Diamond (full 198 questions) TBA
KMMLU TBA
CLIcK (Korean cultural reasoning) TBA
Helmet / Ruler (long context) TBA
Needle-in-Haystack 32 K / 128 K NIAH 32 K: 5/5 per backend (sanity, single model only) — full DUO numbers TBA

Aggregate DUO benchmark results will be published in benchmarks/ after formal evaluation. The verified constituent ranks above are independent third-party measurements and are not aggregate DUO scores.

Cost / latency notes
  • Route mode: comparable to a single 30 B FP8 backend (1× latency, 1× compute).
  • Split mode: ~2× latency (two sequential generations).
  • Ensemble V₁: ~2× compute (parallel) plus a short cross-verify round when majorities disagree.

Bias, Risks, and Limitations
  • Hallucination. Standard LLM caveats apply. Both backends can produce confident but incorrect outputs, especially on out-of-distribution queries.
  • Disagreement bias. Empirically, the V₁ tournament occasionally selects a wrong answer that both backends collectively favor over a single backend's correct one. The gateway exposes the routing decision in _duo_route for auditability.
  • Language coverage. Best performance on English and Korean. Other languages fall back to the closer-fit backend without explicit optimization.
  • Combined weights are not bundled. The aggregator pulls each backend's weights from the constituent repositories. Network and disk for both is required.
  • Two-GPU baseline. BF16 deployment requires two GPUs. FP8 quantization enables single-GPU operation on B200 / H100 (80 GB).
  • Training data cut-off. Darwin-28B-REASON: ~ 2026-Q1. AWAXIS-Think-31B: ~ 2026-Q1.

License

Darwin-60B-DUO inherits the Gemma Terms of Use as its effective combined license — the more restrictive of the two constituent licenses.

Constituent License
Darwin-28B-REASON Apache-2.0
AWAXIS-Think-31B Gemma (inherited from Gemma-4)
Darwin-60B-DUO (aggregate) Gemma (combined-license inheritance)

The orchestrator code ( gateway/ , docker/ ) is offered under Apache-2.0 to maximize developer flexibility; combined-license inheritance applies to served model behavior only.

Please review the Gemma Terms of Use and the Gemma Prohibited Use Policy before commercial deployment.


Acknowledgments
  • FINAL-Bench team — Darwin family architecture and DUO concept
  • Anserwise Korean specialist team — AWAXIS-Think-31B development
  • VIDRAFT — orchestration framework and the Hybrid-A routing strategy
  • Google DeepMind — Gemma-4 foundation
  • Cohere and Qwen teams — Command A+ / Qwen3.5 foundation lineage

Contact

Citation
@misc{darwin60b-duo-2026,
  title  = {Darwin-60B-DUO: A single-endpoint DUO of an English-reasoning SOTA
            and a Korean SOTA via Hybrid-A orchestration},
  author = {FINAL-Bench Team and Anserwise Team and VIDRAFT},
  year   = {2026},
  howpublished = {Hugging Face},
  url    = {https://huggingface.co/FINAL-Bench/Darwin-60B-DUO}
}

Runs of FINAL-Bench Darwin-60B-DUO on huggingface.co

118
Total runs
1
24-hour runs
15
3-day runs
35
7-day runs
86
30-day runs

More Information About Darwin-60B-DUO huggingface.co Model

More Darwin-60B-DUO license Visit here:

https://choosealicense.com/licenses/gemma

Darwin-60B-DUO huggingface.co

Darwin-60B-DUO huggingface.co is an AI model on huggingface.co that provides Darwin-60B-DUO's model effect (), which can be used instantly with this FINAL-Bench Darwin-60B-DUO model. huggingface.co supports a free trial of the Darwin-60B-DUO model, and also provides paid use of the Darwin-60B-DUO. Support call Darwin-60B-DUO model through api, including Node.js, Python, http.

FINAL-Bench Darwin-60B-DUO online free

Darwin-60B-DUO huggingface.co is an online trial and call api platform, which integrates Darwin-60B-DUO's modeling effects, including api services, and provides a free online trial of Darwin-60B-DUO, you can try Darwin-60B-DUO online for free by clicking the link below.

FINAL-Bench Darwin-60B-DUO online free url in huggingface.co:

https://huggingface.co/FINAL-Bench/Darwin-60B-DUO

Darwin-60B-DUO install

Darwin-60B-DUO is an open source model from GitHub that offers a free installation service, and any user can find Darwin-60B-DUO on GitHub to install. At the same time, huggingface.co provides the effect of Darwin-60B-DUO install, users can directly use Darwin-60B-DUO installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Darwin-60B-DUO install url in huggingface.co:

https://huggingface.co/FINAL-Bench/Darwin-60B-DUO

Url of Darwin-60B-DUO

Darwin-60B-DUO huggingface.co Url

Provider of Darwin-60B-DUO huggingface.co

FINAL-Bench
ORGANIZATIONS

Other API from FINAL-Bench