๐
Try it live, no install โ
โ a 35B model answering on a
CPU-only
box.
Pick your build โ
The POCKET lineup โ pick by your device
Repo
File
Size
Runs on
Best for
Korean PPL*
POCKET-35B-GGUF
Q4_K_M
21 GB
PC / server (32 GB RAM)
top quality
5.79
POCKET-35B-GGUF
Q2_K
โญ
13 GB
mini-PC, no GPU
daily driver
6.49
POCKET-35B-GGUF
IQ1_M
8.2 GB
16 GB RAM box
smallest full model
9.69
POCKET-KR-GGUF
IQ2_M
5.1 GB
Android 8 GB+
๐ฐ๐ท Korean phone
7.95
POCKET-KR-MLX
2-bit
5.1 GB
๐
iPhone / iPad / Mac
๐ฐ๐ท Korean, Apple-native
7.95
POCKET-EN-GGUF
iPhone-mix
5.3 GB
๐ iPhone (PocketPal)
๐ English phone
โ
POCKET-EN-GGUF
PC-mix
6.8 GB
PC / Android
๐ English, best quality
โ
*Wikipedia-Korean perplexity, lower is better.
Q4_K_M
= 5.79 baseline. English builds are tuned on English; see each repo.
๐
Why MLX for Korean but GGUF for English on iPhone?
Apple-native MLX only does uniform quantization. Korean survives it (96 experts hold up); English needs our mixed-precision trick, which only GGUF supports โ so the English iPhone build ships as a GGUF you run with
PocketPal
. Honest, not lazy.
Benchmarks โ what is measured, what is not
We measure Bonsai on the same machine with the same stock
llama.cpp
, and we tell you where we lose.
[measured]
Generation speed โ POCKET wins on both CPU and GPU:
POCKET-35B IQ1_M
Bonsai-27B Q1_0
CPU generate (Xeon, 16t)
27.0 tok/s
10.1
๐ข 2.69ร
GPU generate (H100)
197 tok/s
89
๐ข 2.22ร
GPU prompt (H100)
753
1816
๐ด 0.41ร
Quality (HellaSwag, 400q)
61.0%
60.0%
โช tie (CI overlaps)
[measured on a MacBook M3 Pro, 18 GB]
โ and on a laptop, POCKET wins
every
axis, including prompt processing:
POCKET-35B IQ1_M
Bonsai-27B Q1_0
Metal generate (tg64)
25.4 tok/s
12.8
๐ข 1.99ร
CPU generate (8 threads)
13.8 tok/s
4.4
๐ข 3.13ร
Metal prompt (pp128)
240.7 tok/s
73.4
๐ข 3.28ร
CPU prompt (pp128)
45.5 tok/s
9.6
๐ข 4.75ร
On a laptop GPU the arithmetic headroom that let Bonsai win prefill on an H100 is gone, so MoE sparsity wins across the board.
POCKET-35B-Q2_K
runs on the M3 Pro's CPU at
19.5 tok/s
โ on an 18 GB Mac, run Q2_K on CPU (
-ngl 0
); its 13 GB exceeds the recommended Metal budget.
[pending โ community reports welcome]
on-device
iPhone
and
Strix Halo
throughput. We publish only what we ran ourselves; help us fill the rest.
The same-size rival
Ternary-Bonsai-27B-Q2_0
(7.2 GB)
fails to load in upstream llama.cpp
โ it needs the PrismML fork. POCKET runs on the tools you already have.
Files in this repo
File
Size
Experts
Runs on
Korean PPL
POCKET-KR-IQ2_M.gguf
โญ
5.1 GB
96
Android 8 GB+
7.95
POCKET-KR-160-Q2_K.gguf
8.5 GB
160
phone/PC
6.88 (quality-first)
POCKET-KR-160-Q3_K_M.gguf
11 GB
160
PC
6.35
Pruned from 256 experts to the ones Korean actually uses (94% routing coverage at 128). Active params unchanged โ
same speed, half the size.
POCKET is quantized from
Darwin-36B-Opus
, VIDRAFT's flagship โ a model bred and evolved over several generations on the
Darwin platform
(crossbreeding, healing, expert surgery). Darwin-36B-Opus itself traces back to a Qwen3.5-family MoE architecture.
The CPU/GPU speed comes from the sparse-MoE architecture plus ordinary quantization โ reproducible with the same base and the same tools. What we add is the Darwin-evolved weights, the honest measurement, the Korean tuning, and the pruning that makes the 5 GB phone builds.
Limitations
The iPhone/Mac speed is
not yet measured by us
โ community reports welcome.
Extreme quants (
IQ1_M
) hurt Korean ~2.8ร more than English; use
Q2_K
or larger for quality.
English phone builds trade quality for size; the PC build (
PC-mix
) is much closer to full quality.
License
Apache-2.0.
POCKET is a VIDRAFT model family. 35B, in your pocket. No GPU.
Runs of FINAL-Bench POCKET-KR-GGUF on huggingface.co
995
Total runs
7
24-hour runs
8
3-day runs
477
7-day runs
808
30-day runs
More Information About POCKET-KR-GGUF huggingface.co Model
POCKET-KR-GGUF huggingface.co is an AI model on huggingface.co that provides POCKET-KR-GGUF's model effect (), which can be used instantly with this FINAL-Bench POCKET-KR-GGUF model. huggingface.co supports a free trial of the POCKET-KR-GGUF model, and also provides paid use of the POCKET-KR-GGUF. Support call POCKET-KR-GGUF model through api, including Node.js, Python, http.
POCKET-KR-GGUF huggingface.co is an online trial and call api platform, which integrates POCKET-KR-GGUF's modeling effects, including api services, and provides a free online trial of POCKET-KR-GGUF, you can try POCKET-KR-GGUF online for free by clicking the link below.
FINAL-Bench POCKET-KR-GGUF online free url in huggingface.co:
POCKET-KR-GGUF is an open source model from GitHub that offers a free installation service, and any user can find POCKET-KR-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of POCKET-KR-GGUF install, users can directly use POCKET-KR-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.