๐ฃ Tiny footprint, big brain โ a local
coding
model for
everyone
No matter your GPU. No matter your RAM.
If you've got
~4.5 GB
of VRAM
or
unified memory free,
you can run your own private, offline coding assistant right now. ๐
This is the
v1 / code edition
โ distilled from
real chain-of-thought
so it
thinks through
a problem
before writing the solution. ๐ง ๐ป All local, all yours, no API, no cloud.
๐ฏ What it is
A focused fine-tune of Gemma 4 12B on
verifiable Python coding
data โ every training example's reasoning leads to
code that
actually passed its tests
. The result reasons in the open (edge cases, complexity, approach) and then
emits a clean, runnable solution. ๐
๐ Announcements
๐๐ฅ IT'S HERE โ v2 is OUT NOW!
v2 has shipped โ the
GGUF quants are live and ready to run
โ
grab v2 here
. ๐
The
full
safetensors
master
(build / fine-tune on top) goes up
tomorrow
. v2 is
agentic + coding
focused โ
the piece v1 was missing.
Here's the result that got me most excited.
When I saw v2's
tau2-bench
telecom
result โ an agentic tool-use
benchmark where the model has to
diagnose โ fix โ verify
, exactly like real terminal/debugging work โ I literally got
launched out of my chair
(โฆokay,
kidding
๐). The jump in
actually solving the problem
is wild:
tau2-bench
telecom
ยท local, same harness,
Q8_0
score
official
gemma-4-12B-it
(base)
~15%
๐ข
v2 (this release)
~55%
The base model tends to
give up early
(hands the problem off to a human);
v2 keeps going
and works it the way a
much bigger model would. Full benchmark details are in the
v2 card
now. ๐ง
โ safetensors master (this v1 model) is UP.
Full-precision weights are live โ
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1
โ roll your own GGUF / MLX / AWQ quants or fine-tune straight from the master. ๐
A community member spotted that this model was reporting only a
131K
context window. That turned out to be
the well-known upstream
Gemma 4 metadata bug
โ Google's initial
config.json
shipped with
max_position_embeddings: 131072
instead of the real
262144 (256K)
, and that value got baked into a lot of
downstream finetunes and quants (including this one) before it was fixed upstream.
The weights were always fine โ it was purely a metadata field.
All GGUF quants have been re-patched to the
full 256K context
(
gemma4.context_length = 262144
). Just re-download if you grabbed an earlier copy. ๐
๐ Training data (the interesting part ๐ณ)
This is a
distillation
of two complementary chain-of-thought sources, both over verifiable Python coding tasks
(algorithmic / function-level problems that come with deterministic tests):
๐ฅ Main set โ Composer 2.5
real
CoT.
Genuine, model-authored reasoning traces. The teacher solved each problem,
its code was
run against the task's tests, and only the passing solutions were kept
. So the reasoning you're
learning from leads to code that
actually works
.
๐ฅ Aux set โ Fable 5 (released today! ๐).
A clever twist: we took the problems where
Composer 2.5 got it wrong
and handed them to
Fable 5
to
redo
โ re-deriving a fresh, self-consistent chain-of-thought and a correct
solution, again
gated on passing the tests
. This recovers the hard cases the main teacher missed. These traces
are
synthetic
(rationalized CoT), and are tagged separately so the two sources stay distinguishable.
The recipe: real CoT for the bulk of solid coverage, plus synthetic "second-attempt" CoT to patch the failures โ
both verified by execution before anything entered training. โ
๐ฆ Pick your size (GGUF quants)
Quant
Size
Vibe
๐ข
Q2_K
4.5 GB
tiniest โ runs almost anywhere
๐ก
Q3_K_M
5.7 GB
great for 8 GB VRAM โ much better than Q2
๐ต
Q4_K_M
6.87 GB
the sweet spot ๐ (recommended)
๐ฃ
Q6_K
9.11 GB
near-lossless
โช
Q8_0
11.8 GB
basically full quality
๐งฎ "Will it fit?" โ context length cheat-sheet
Rough estimates ๐ค (assumes
q8_0
KV cache + ~1.5 GB overhead;
use
q4_0
KV cache for โ2ร more context!
).
Max context is
256K
. "โ" = won't fit, pick a smaller quant. โ๏ธ
Your VRAM / unified mem
๐ข Q2_K (4.5G)
๐ก Q3_K_M (5.7G)
๐ต Q4_K_M (6.87G)
๐ฃ Q6_K (9.11G)
โช Q8_0 (11.8G)
8 GB
~16K ctx
~10K
tight (~2โ4K)
โ
โ
12 GB
~48K
~38K
~30K
~12K
โ
16 GB
~80K
~72K
~64K
~44K
~22K
24 GB
~200K
~160K
~128K
~110K
~88K
32 GB
256K (max) ๐
256K
256K
~230K
~190K
๐ก Apple Silicon / integrated GPUs with
unified memory
count too โ same numbers, just slower than a dGPU.
๐ก Low on room? Drop a quant or switch KV cache to
q4_0
and your context roughly doubles.
๐ How to run it (super easy)
Option A โ llama.cpp (recommended) ๐ฆ
Grab a quant above (e.g.
โฆ-Q4_K_M.gguf
) and
llama-server
from
llama.cpp
.
โ ๏ธ Needs a
recent llama.cpp
(this is the
gemma4_unified
architecture โ older builds won't load it).
Run a server (Windows
.bat
shown โ tweak
--port
,
--ctx-size
to taste):
Open
http://localhost:18080
and chat. ๐ (Tip: bump
--ctx-size
per the table; use
q4_0
KV for more.)
Option B โ one-click apps ๐ฑ๏ธ
Works in
LM Studio
,
Jan
,
Ollama
, etc. โ just import the GGUF, pick your quant, go. ๐พ
๐ง Thinking mode
This model thinks in Gemma's native thought channel before answering โ exactly how it was trained. Keep
enable_thinking=true
(the default chat template handles it). Recommended sampling:
temp 1.0, top_p 0.95, top_k 64
.
For coding you can also go greedy (
temp 0
) for more deterministic solutions.
โ ๏ธ Good to know
Reduced refusals:
the training data is task-focused with no safety hedging, so this refuses less than the base
model. It is
not
safety-aligned โ add your own guardrails for production. Use responsibly. ๐
Specialized for
Python / algorithmic
coding. Reasoning quality is strongest in that domain; general-knowledge
facts/numbers should still be double-checked.
English-centric.
๐ Base & License
License: Apache 2.0.
Gemma 4 is released by Google under
Apache 2.0
(unlike the older Gemma 1/2/3 terms), so this fine-tune is
Apache 2.0
too โ free to use, modify, and redistribute. ๐
gemma-4-12b-coder-fable5-composer2.5-8bit huggingface.co is an AI model on huggingface.co that provides gemma-4-12b-coder-fable5-composer2.5-8bit's model effect (), which can be used instantly with this mlx-community gemma-4-12b-coder-fable5-composer2.5-8bit model. huggingface.co supports a free trial of the gemma-4-12b-coder-fable5-composer2.5-8bit model, and also provides paid use of the gemma-4-12b-coder-fable5-composer2.5-8bit. Support call gemma-4-12b-coder-fable5-composer2.5-8bit model through api, including Node.js, Python, http.
gemma-4-12b-coder-fable5-composer2.5-8bit huggingface.co is an online trial and call api platform, which integrates gemma-4-12b-coder-fable5-composer2.5-8bit's modeling effects, including api services, and provides a free online trial of gemma-4-12b-coder-fable5-composer2.5-8bit, you can try gemma-4-12b-coder-fable5-composer2.5-8bit online for free by clicking the link below.
mlx-community gemma-4-12b-coder-fable5-composer2.5-8bit online free url in huggingface.co:
gemma-4-12b-coder-fable5-composer2.5-8bit is an open source model from GitHub that offers a free installation service, and any user can find gemma-4-12b-coder-fable5-composer2.5-8bit on GitHub to install. At the same time, huggingface.co provides the effect of gemma-4-12b-coder-fable5-composer2.5-8bit install, users can directly use gemma-4-12b-coder-fable5-composer2.5-8bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
gemma-4-12b-coder-fable5-composer2.5-8bit install url in huggingface.co: