Gemma 4 26B A4B
, self-quantized to MLX by
Atomic Chat
. Built straight from Google's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
Highlights
25.2B total / 3.8B active per token parameters
: the weights this repo quantizes.
Context length
: 256K tokens, as published by Google.
30 layers
: Mixture-of-Experts, hybrid sliding-window (1024) and global attention.
Modalities
: Text, Image.
Full imatrix ladder
: every quant is calibrated with an importance matrix.
Reasoning
: All models in the family are designed as highly capable reasoners, with configurable thinking modes.
Diverse & Efficient Architectures
: Offers Dense and Mixture-of-Experts (MoE) variants of different sizes for scalable deployment.
These MLXs are
self-quantized from the original weights
, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Model Overview
Property
Value
Base model
google/gemma-4-26B-A4B-it
Parameters
25.2B total / 3.8B active per token
Layers
30
Experts
128 routed (top-8)
Sliding window
1024 tokens
Context length
256K tokens
Vocabulary
262K
Modalities
Text, Image
Architecture
Mixture-of-Experts, 128 experts (top-8), hybrid sliding-window (1024) and global attention, 16 attention heads over 8 KV heads,
Gemma4ForConditionalGeneration
This repo
MLX weights
Benchmarks
Benchmark
Score
MMLU Pro
82.6%
AIME 2026 no tools
88.3%
LiveCodeBench v6
77.1%
Codeforces ELO
1718
GPQA Diamond
82.3%
Tau2 (average over 3)
68.2%
HLE no tools
8.7%
HLE with search
17.2%
BigBench Extra Hard
64.8%
MMMLU
86.3%
MMMU Pro
73.8%
OmniDocBench 1.5 (average edit distance, lower is better)
0.149
MATH-Vision
82.4%
MedXPertQA MM
58.1%
MRCR v2 8 needle 128k (average)
44.1%
Scores are Google's published results for the base
google/gemma-4-26B-A4B-it
, not our own measurements. Quantization preserves the large majority of this;
Q4_K_M
and up stay close to full precision.
Get started
Atomic Chat
:
search
AtomicChat/gemma-4-26B-A4B-it-MLX-4bit
and hit
Use this model
.
gemma-4-26B-A4B-it-MLX-4bit huggingface.co is an AI model on huggingface.co that provides gemma-4-26B-A4B-it-MLX-4bit's model effect (), which can be used instantly with this AtomicChat gemma-4-26B-A4B-it-MLX-4bit model. huggingface.co supports a free trial of the gemma-4-26B-A4B-it-MLX-4bit model, and also provides paid use of the gemma-4-26B-A4B-it-MLX-4bit. Support call gemma-4-26B-A4B-it-MLX-4bit model through api, including Node.js, Python, http.
gemma-4-26B-A4B-it-MLX-4bit huggingface.co is an online trial and call api platform, which integrates gemma-4-26B-A4B-it-MLX-4bit's modeling effects, including api services, and provides a free online trial of gemma-4-26B-A4B-it-MLX-4bit, you can try gemma-4-26B-A4B-it-MLX-4bit online for free by clicking the link below.
AtomicChat gemma-4-26B-A4B-it-MLX-4bit online free url in huggingface.co:
gemma-4-26B-A4B-it-MLX-4bit is an open source model from GitHub that offers a free installation service, and any user can find gemma-4-26B-A4B-it-MLX-4bit on GitHub to install. At the same time, huggingface.co provides the effect of gemma-4-26B-A4B-it-MLX-4bit install, users can directly use gemma-4-26B-A4B-it-MLX-4bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
gemma-4-26B-A4B-it-MLX-4bit install url in huggingface.co: