Grug 12B is a compact-reasoning fine-tune of
google/gemma-4-12B-it
.
It was trained to keep the useful information from a reasoning trace while
making the trace shorter, more direct, and less polished.
This repository is published as merged Transformers/safetensors model weights.
It was trained with QLoRA, then merged into the base model for release.
GGUF
A llama.cpp
Q4_K_M
GGUF release is available in the adjacent repo:
kai-os/Grug-12B-GGUF
.
What Changed
The training target is a terse internal-reasoning style: short high-density
steps, fewer filler words, and explicit preservation of key constraints,
branching decisions, invariants, edge cases, and final-answer checks.
The goal is lower reasoning-token usage relative to the base model while
preserving answer quality. It is not meant to hide uncertainty or remove needed
reasoning.
Training Data
The data pipeline started from a recent, filtered reasoning pool and converted
verbose traces into compact traces before SFT packing.
Source gate:
Run date: June 30, 2026.
Default freshness cutoff: 45 days. Sources older than May 16, 2026 were
rejected unless manually allowed.
The compact reasoning transform was generated with
cyankiwi/Qwen3.6-35B-A3B-AWQ-4bit
served by vLLM. Rows were checked for
compression ratio, answer preservation, and obvious loss of critical reasoning
information before training.
Training Procedure
Training was completion-only SFT: prompt tokens were masked with
-100
, and
only the assistant completion was trained.
Core settings:
Base model:
google/gemma-4-12B-it
.
Method: QLoRA / PEFT LoRA, merged into full model weights for upload.
Quantization during training: 4-bit NF4 with BF16 compute.
Train runtime: about 35 minutes 20 seconds on one A100.
Final eval loss: 0.8895.
No train or validation rows were skipped in the final run.
Local Evaluation
Small local EOS-only math proxy eval, no generation token cap:
Model
Rows
Total generated tokens
Avg generated tokens
Proxy accuracy
Numeric last-match rate
google/gemma-4-12B-it
base
36
8,227
228.53
91.7%
86.1%
Grug 12B
36
2,482
68.94
100.0%
100.0%
This is a small proxy eval, not a broad benchmark. Treat it as a smoke test
showing the intended token-efficiency direction, then run your own benchmark.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "kai-os/Grug-12B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
model.eval()
messages = [
{"role": "user", "content": "If a shirt is $80 and goes 25% off, what is the sale price?"}
]
inputs = tokenizer.apply_chat_template(
messages,
return_tensors="pt",
add_generation_prompt=True,
).to(model.device)
with torch.no_grad():
output = model.generate(inputs, do_sample=False, max_new_tokens=512)
print(tokenizer.decode(output[0], skip_special_tokens=True))
For token-efficiency tests, compare against the base model with the same prompt,
same decoding settings, and no artificial token cap unless your deployment
requires one.
Limitations
This is an experimental fine-tune.
It may over-compress reasoning on tasks that need longer derivations.
It inherits the base model's limitations and safety behavior.
The reported eval is small and local.
The dataset includes synthetic and distilled reasoning traces from the
listed open datasets; review source licenses and provenance before using this
in commercial or sensitive settings.
Acknowledgements
Thanks to
Lambda
, the inference provider, for compute
credits that supported the dataset work, training, and evaluation.
Runs of kai-os Grug-12B on huggingface.co
136
Total runs
0
24-hour runs
-17
3-day runs
-10
7-day runs
-250
30-day runs
More Information About Grug-12B huggingface.co Model
Grug-12B huggingface.co is an AI model on huggingface.co that provides Grug-12B's model effect (), which can be used instantly with this kai-os Grug-12B model. huggingface.co supports a free trial of the Grug-12B model, and also provides paid use of the Grug-12B. Support call Grug-12B model through api, including Node.js, Python, http.
Grug-12B huggingface.co is an online trial and call api platform, which integrates Grug-12B's modeling effects, including api services, and provides a free online trial of Grug-12B, you can try Grug-12B online for free by clicking the link below.
kai-os Grug-12B online free url in huggingface.co:
Grug-12B is an open source model from GitHub that offers a free installation service, and any user can find Grug-12B on GitHub to install. At the same time, huggingface.co provides the effect of Grug-12B install, users can directly use Grug-12B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.