Darwin-Evolved on Qwen3.6-35B-A3B | 35B total / ~3B active | ๐ฐ๐ท Korean-Specialized | Thinking Mode | Hybrid Linear/Full Attention | Multi-Token Prediction | 262K Context | BF16 | Apache 2.0
Darwin FFN-level evolutionary merge โ Korean specialization โ 86.36% on GPQA Diamond (majority-of-8+)
Abstract
Ourbox-35B-JGOS
is a 35-billion-parameter mixture-of-experts (MoE) reasoning model produced by the
Darwin
evolutionary breeding platform (FINAL-Bench / VIDRAFT_LAB). Rather than retraining from scratch, Darwin recombines the
feed-forward (FFN / MoE expert) tensors
of the
Qwen3.6-35B-A3B
backbone with those of additional specialized donor models, then
evolves
the merged descendant toward a target objective โ here,
Korean-language specialization
.
Because the merge operates at the
expert-FFN level
, Ourbox inherits complementary domain and language competencies from multiple sources while preserving the backbone's hybrid-attention topology and 262K long-context behavior. The result is a Korean-specialized reasoning model that remains highly competitive on English graduate-level science: on
GPQA Diamond
(198 questions across physics, chemistry, biology), Ourbox-35B-JGOS scores
86.36% (171/198)
under a majority-of-8+ protocol. On Hugging Face's live GPQA leaderboard this
improves on its own Qwen3.6-35B-A3B backbone (86.0)
by +0.36 points and edges past
GLM-5.1 (86.2)
and
GLM-5 (86.0)
โ with only
~3B active parameters
.
GPQA Diamond Leaderboard โ Hugging Face
Idavidrein/gpqa
(2026-07-11)
Ourbox-35B-JGOS on the
official Hugging Face GPQA Diamond leaderboard
(
Idavidrein/gpqa
, base-model view, 50 models). FINAL-Bench models in
bold
:
#
Model
GPQA Diamond
1
zai-org/GLM-5.2
91.2
2
FINAL-Bench/Darwin-398B-JGOS
90.9
3
moonshotai/Kimi-K2.6
90.5
4
tencent/Hy3
90.4
5
deepseek-ai/DeepSeek-V4-Pro
90.1
6
FINAL-Bench/Darwin-28B-REASON
89.39
7
Qwen/Qwen3.5-397B-A17B
88.4
8
FINAL-Bench/Darwin-36B-Opus
88.4
9
FINAL-Bench/Darwin-60B-DUO
88.38
10
inclusionAI/Ring-2.6-1T
88.27
11
deepseek-ai/DeepSeek-V4-Flash
88.1
12
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B (NVFP4)
87.9
13
zai-org/GLM-4.7-FP8
87.88
14
Qwen/Qwen3.6-27B
87.8
15
moonshotai/Kimi-K2.5
87.6
16
moonshotai/Kimi-K2.5
(source)
87.37
17
tencent/Hy3-preview
87.2
18
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B (BF16)
87.0
19
FINAL-Bench/Darwin-27B-Opus
86.9
20
Qwen/Qwen3.5-122B-A10B
86.6
โ 21
FINAL-Bench/Ourbox-35B-JGOS
๐ฐ๐ท
86.36
22
zai-org/GLM-5.1
86.2
23
zai-org/GLM-5
86.0
24
Qwen/Qwen3.6-35B-A3B
(Ourbox backbone)
86.0
25
FINAL-Bench/Darwin-31B-Opus
85.9
The FINAL-Bench Darwin family dominates the upper board โ
5 of the 20 models ranked above Ourbox are Darwin models
(Darwin-398B-JGOS #2, Darwin-28B-REASON #6, Darwin-36B-Opus #8, Darwin-60B-DUO #9, Darwin-27B-Opus #19). At
86.36%
, Ourbox-35B-JGOS ranks
#21 of 50
on the live leaderboard and โ most notably โ
improves on its own Qwen3.6-35B-A3B backbone (86.0, #24) by +0.36 points
, confirming that the Darwin FFN-merge and Korean specialization
added
capability rather than eroding it. It also edges past
GLM-5.1 (86.2, #22)
and
GLM-5 (86.0, #23)
while activating only ~3B parameters.
Ranks reflect the live leaderboard as of 2026-07-11 (which counts quantized/duplicate entries); positions shift as it updates. Ourbox-35B-JGOS is
live and listed at #21
.
What Is Darwin?
Darwin
is the evolutionary model-breeding platform developed by FINAL-Bench / VIDRAFT_LAB. Rather than allocating further compute to gradient optimization, Darwin treats trained checkpoints as a
genetic pool
and discovers high-performing descendants through principled recombination of their weight tensors โ with a particular focus on the
FFN / MoE expert
subspace, where domain and language competence is concentrated.
At a high level, the platform performs:
Per-tensor compatibility analysis
across the backbone and donor models to identify which FFN experts and components transfer cleanly and which require weighted recombination.
FFN-level merge & evolution
โ the descendant's expert tensors are assembled from the pool and iteratively evolved toward a target objective (Korean specialization for Ourbox).
Verification
via a multi-phase scientific benchmark before release.
Specific algorithmic details of the Darwin engine are proprietary to FINAL-Bench. All Darwin models are released under the base model's open-source license (Apache 2.0).
JGOS
is the reasoning-model line built with Darwin;
Ourbox
is its Korean-specialized 35B-A3B member.
Evolution Process
Ourbox-35B-JGOS is bred, not trained:
Backbone
:
Qwen/Qwen3.6-35B-A3B
โ the foundation MoE, contributing its hybrid-attention topology (ยพ linear + ยผ full), 256-expert routing, MTP head, and 262K context.
FFN donors
: additional specialized models whose
feed-forward / expert tensors
are recombined into the backbone by the Darwin engine, contributing complementary domain and Korean-language competence.
Evolution objective
: Korean specialization โ the evolutionary selection biases the merged expert population toward stronger Korean reasoning and generation, while structural and long-context behavior is inherited intact from the backbone.
The merge operates
without gradient optimization on the final assembly
; a deployable bfloat16 checkpoint is produced by the Darwin pipeline directly.
๐ฐ๐ท Korean Specialization
Ourbox-35B-JGOS is specialized for
Korean
. The Darwin evolutionary process selects and recombines FFN experts to strengthen Korean-language reasoning, comprehension, and generation โ targeting natural Korean output, robust handling of Korean scientific/technical text, and reduced character-level corruption on large Korean inputs.
Crucially, this specialization does
not
come at the cost of general capability: the model's
86.36% GPQA Diamond
(in English) improves on its own Qwen3.6-35B-A3B backbone (86.0) on Hugging Face's live leaderboard, evidence that the FFN-merge preserved scientific reasoning depth while adding Korean strength. The model remains fully multilingual (Korean-first), with English, Chinese, Japanese, and other languages inherited from the backbone.
Architecture
Ourbox-35B-JGOS retains the full Qwen3.6-35B-A3B architecture (
qwen3_5_moe
codebase):
The
hybrid attention
design (ยพ linear + ยผ full) gives near-linear KV-cache scaling across the 262K window, and the
Multi-Token Prediction
head provides a built-in draft for speculative decoding.
GPQA Diamond Evaluation
Methodology
Ourbox-35B-JGOS was evaluated on all
198 GPQA Diamond
questions using a two-pass
majority-of-8+
protocol (identical to sibling FINAL-Bench reasoning models, for cross-model comparability):
Pass 1 โ Greedy baseline
All 198 questions, deterministic decoding (
do_sample=False
)
Up to 5,120 new tokens per question (full
<think>
trajectories)
Standard multiple-choice prompt format
Pass 2 โ Stochastic majority vote with tiebreaker
Each question is answered by
8 independent stochastic generations
(
temperature=0.7
,
max_tokens=5120
); the majority answer is taken
Where the 8-vote margin is inconclusive (e.g. 3:3 / 3:4 / 4:4), an additional
16-vote tiebreaker
round (
temperature=0.5
) resolves the answer
The final answer for each question is extracted after the
</think>
delimiter.
Result
Metric
Value
Correct
171 / 198
GPQA Diamond accuracy (maj@8+)
86.36%
Evaluated against the
Idavidrein/gpqa
gpqa_diamond
split. The majority-of-8+ protocol surfaces answers that greedy decoding leaves subdominant โ a pattern characteristic of well-formed chain-of-thought models โ carrying Ourbox above its Qwen3.6-35B-A3B backbone (86.0) and past GLM-5.1 (86.2) on graduate-level science.
Temperature
: 0.6โ0.7 for reasoning / majority voting; 0.0 for greedy deterministic
max_new_tokens
: โฅ5120 to accommodate full
<think>
trajectories
Chat template
: assistant turn opens with
<think>
when
apply_chat_template(add_generation_prompt=True)
is used
VRAM Requirements
Precision
VRAM
Recommended GPU
bf16 (full)
~72 GB
1ร H100 80GB / 1ร B200
8-bit
~40 GB
1ร A100 40GB+ / 1ร L40S
4-bit
~22 GB
1ร RTX 4090 / 1ร A10
Key Findings
Korean specialization without capability loss.
Darwin's FFN-level merge adds Korean-language strength while retaining
86.36% GPQA Diamond
โ above the model's own Qwen3.6-35B-A3B backbone (86.0). Specialization and general reasoning are not a zero-sum trade under expert-level recombination.
Specialization improves on the backbone.
On Hugging Face's live GPQA Diamond leaderboard, Ourbox (86.36) exceeds its own Qwen3.6-35B-A3B backbone (86.0) and edges past GLM-5.1 (86.2) and GLM-5 (86.0) โ the Darwin FFN-merge added Korean capability without eroding scientific reasoning, at ~3B active parameters.
Breeding beats retraining for specialization.
A deployable, Korean-specialized 35B checkpoint is produced by evolutionary FFN recombination โ no full-model gradient training on the final assembly โ demonstrating Darwin as an efficient route to targeted, high-capability models.
References
Rein et al.,
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
, 2024.
dataset
Qwen Team,
Qwen3.6 Technical Report
, 2026.
Built By
FINAL-Bench / VIDRAFT_LAB
โ Darwin evolutionary breeding platform, JGOS Korean-specialized reasoning line.
Backbone weights by the Qwen Team (Qwen3.6-35B-A3B). Released under Apache 2.0.
Citation
@misc{ourbox-35b-jgos,
title = {Ourbox-35B-JGOS: Korean-Specialized, Darwin-Evolved 35B-A3B Reasoning MoE},
author = {FINAL-Bench and VIDRAFT_LAB},
year = {2026},
url = {https://huggingface.co/FINAL-Bench/Ourbox-35B-JGOS},
note = {Qwen3.6-35B-A3B backbone, Darwin FFN-level evolutionary merge, Korean-specialized, 86.36% GPQA Diamond (maj@8+)}
}
Runs of FINAL-Bench Ourbox-35B-JGOS on huggingface.co
47
Total runs
0
24-hour runs
-5
3-day runs
-4
7-day runs
-143
30-day runs
More Information About Ourbox-35B-JGOS huggingface.co Model
Ourbox-35B-JGOS huggingface.co is an AI model on huggingface.co that provides Ourbox-35B-JGOS's model effect (), which can be used instantly with this FINAL-Bench Ourbox-35B-JGOS model. huggingface.co supports a free trial of the Ourbox-35B-JGOS model, and also provides paid use of the Ourbox-35B-JGOS. Support call Ourbox-35B-JGOS model through api, including Node.js, Python, http.
Ourbox-35B-JGOS huggingface.co is an online trial and call api platform, which integrates Ourbox-35B-JGOS's modeling effects, including api services, and provides a free online trial of Ourbox-35B-JGOS, you can try Ourbox-35B-JGOS online for free by clicking the link below.
FINAL-Bench Ourbox-35B-JGOS online free url in huggingface.co:
Ourbox-35B-JGOS is an open source model from GitHub that offers a free installation service, and any user can find Ourbox-35B-JGOS on GitHub to install. At the same time, huggingface.co provides the effect of Ourbox-35B-JGOS install, users can directly use Ourbox-35B-JGOS installed effect in huggingface.co for debugging and trial. It also supports api for free installation.