MLX q8
— the
native training/serving precision
for Patak (SFT and DPO were both trained directly
on top of a q8-quantized base, not fused-then-quantized after the fact). See the
patak/
repo's README
for full architecture, CPT/SFT/DPO training details, and benchmarks — this file covers only the
q8-specific notes.
Quantization
q8, group size 64
Size on disk
~9.1 GB (vs. ~17 GB bf16)
Quality
This
is
the model's native precision — the
patak/
bf16 repo is dequantized
from
this, not the other way around. Own benchmark testing found bf16 and q8 score within 1 point of each other (218 vs. 217/250).
Max context length
32,768 tokens (EuroLLM-9B's native context)
Usage
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tok = load("emese-tech/patak-mlx")
p = tok.apply_chat_template([{"role": "user", "content": "Mi Magyarország fővárosa?"}],
tokenize=False, add_generation_prompt=True)
print(generate(model, tok, prompt=p, max_tokens=256, sampler=make_sampler(temp=0.2)))
Decode:
temperature
0.2
, no repetition penalty, eos
{2, 4}
, ChatML template.
⚠️
This repo is
mlx_lm
-only
— MLX's q8 quantization packs weights into
uint32
+ per-group
scales
/
biases
tensors with a
quantization
block in
config.json
that plain
transformers
does not
understand. Use the
patak/
(bf16) repo for
transformers
/vLLM/TGI.
Training
This is the primary artifact of the SFT+DPO training chain — see
patak/README.md
for the full CPT
(5.1M tokens/5,000 iters), SFT (
instruct_v18b
, 1 epoch, rank16/scale32/lr1.5e-5), and DPO (36 alfa
pairs, 120 iters, rank16/scale32/lr5e-6) recipe, plus the 218/250 Ultimate · 302/376 BlindSpot benchmark
results.
Benchmarks
This exact q8 artifact scored 413/500 (83%) on emese-bench v1
(the consolidated 500-pt
Ultimate+BlindSpot benchmark) — by far the family's strongest result, near-perfect on longform,
reading, code, safety, honesty, and English. See
emese-bench/results/patak-mlx.md
for the full
category-by-category transcript and
emese-bench/README.md
for the benchmark's design.
Runs of emese-tech patak-mlx on huggingface.co
202
Total runs
1
24-hour runs
3
3-day runs
6
7-day runs
202
30-day runs
More Information About patak-mlx huggingface.co Model
patak-mlx huggingface.co is an AI model on huggingface.co that provides patak-mlx's model effect (), which can be used instantly with this emese-tech patak-mlx model. huggingface.co supports a free trial of the patak-mlx model, and also provides paid use of the patak-mlx. Support call patak-mlx model through api, including Node.js, Python, http.
patak-mlx huggingface.co is an online trial and call api platform, which integrates patak-mlx's modeling effects, including api services, and provides a free online trial of patak-mlx, you can try patak-mlx online for free by clicking the link below.
emese-tech patak-mlx online free url in huggingface.co:
patak-mlx is an open source model from GitHub that offers a free installation service, and any user can find patak-mlx on GitHub to install. At the same time, huggingface.co provides the effect of patak-mlx install, users can directly use patak-mlx installed effect in huggingface.co for debugging and trial. It also supports api for free installation.