Homunculus
is a 12 billion-parameter instruction model distilled from
Qwen3-235B
onto the
Mistral-Nemo
backbone.
It was purpose-built to preserve Qwen’s two-mode interaction style—
/think
(deliberate chain-of-thought) and
/nothink
(concise answers)—while running on a single consumer GPU.
✨ What’s special?
Feature
Detail
Reasoning-trace transfer
Instead of copying just final probabilities, we align
full
logit trajectories, yielding more faithful reasoning.
Total-Variation-Distance loss
To better match the teacher’s confidence distribution and smooth the loss landscape.
Tokenizer replacement
The original Mistral tokenizer was swapped for Qwen3's tokenizer.
Dual interaction modes
Use
/think
when you want transparent step-by-step reasoning (good for analysis & debugging). Use
/nothink
for terse, production-ready answers. Most reliable in the system role field.
Benchmark results
Benchmark
Score
GPQADiamond (average of 3)
57.1%
mmlu
67.5%
🔧 Quick Start
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "arcee-ai/Homunculus"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto"
)
# /think mode - Chain-of-thought reasoning
messages = [
{"role": "system", "content": "You are a helpful assistant. /think"},
{"role": "user", "content": "Why is the sky blue?"},
]
output = model.generate(
tokenizer.apply_chat_template(messages, tokenize=True, return_tensors="pt"),
max_new_tokens=512,
temperature=0.7
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
# /nothink mode - Direct answers
messages = [
{"role": "system", "content": "You are a helpful assistant. /nothink"},
{"role": "user", "content": "Summarize the plot of Hamlet in two sentences."},
]
output = model.generate(
tokenizer.apply_chat_template(messages, tokenize=True, return_tensors="pt"),
max_new_tokens=128,
temperature=0.7
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
💡 Intended Use & Limitations
Homunculus is designed for:
Research
on reasoning-trace distillation, Logit Imitation, and mode-switchable assistants.
Lightweight production
deployments that need strong reasoning at <12 GB VRAM.
Known limitations
May inherit biases from the Qwen3 teacher and internet-scale pretraining data.
Long-context (>32 k tokens) use is experimental—expect latency & memory overhead.
Runs of arcee-ai Homunculus on huggingface.co
301
Total runs
0
24-hour runs
30
3-day runs
36
7-day runs
164
30-day runs
More Information About Homunculus huggingface.co Model
Homunculus huggingface.co is an AI model on huggingface.co that provides Homunculus's model effect (), which can be used instantly with this arcee-ai Homunculus model. huggingface.co supports a free trial of the Homunculus model, and also provides paid use of the Homunculus. Support call Homunculus model through api, including Node.js, Python, http.
Homunculus huggingface.co is an online trial and call api platform, which integrates Homunculus's modeling effects, including api services, and provides a free online trial of Homunculus, you can try Homunculus online for free by clicking the link below.
arcee-ai Homunculus online free url in huggingface.co:
Homunculus is an open source model from GitHub that offers a free installation service, and any user can find Homunculus on GitHub to install. At the same time, huggingface.co provides the effect of Homunculus install, users can directly use Homunculus installed effect in huggingface.co for debugging and trial. It also supports api for free installation.