Archon
is a fine-tuned, locally-deployable AI agent built on Google's
Gemma 4 E4B
architecture. It is designed for autonomous, long-horizon reasoning and agentic task execution on consumer hardware — specifically targeting deployment on 8GB VRAM GPUs via the Q4_K_M GGUF format.
Archon is sharp, slightly edgy, deeply sarcastic, and flawlessly effective.
Model Details
Property
Value
Base Model
google/gemma-4-E4B-it
Architecture
Dense Transformer with Per-Layer Embeddings (PLE)
Total Parameters
~8B (4.5B effective)
Fine-Tuning Method
LoRA (rank 64, rsLoRA, α=64)
Training Framework
Unsloth 2026.4.2 + TRL 0.24
Context Length
4096 tokens (training) / 128k (inference)
Training Hardware
NVIDIA A100 SXM4-80GB
Quantization (GGUF)
Q4_K_M (~4.1–4.5GB, optimal for 8GB VRAM)
Developer
DuoNeural
License
Gemma Terms of Use
Intended Use
Archon is designed for:
Local agentic deployment
on hardware with 8GB VRAM (GTX 1070, RTX 3060, etc.)
Long-horizon autonomous reasoning
with
<think>
chain-of-thought
Tool use and terminal execution
via Hermes Agent or OpenClaw
Complex coding and software architecture
tasks
Persistent, always-on AI assistant with accumulated skill memory
# Pull and run locally
ollama pull DuoNeural/Archon-Gemma-4-E4B-Q4_K_M
ollama run DuoNeural/Archon-Gemma-4-E4B-Q4_K_M
Measured Inference Performance
NVIDIA GTX 1070 (8GB VRAM, Pascal) — Q4_K_M via Ollama
Metric
Value
Generation speed
32.90 tokens/s
Prefill speed
175.92 tokens/s
Model load time
394 ms
VRAM used
~5.0 GB
This is remarkably fast for an 8B-class model on a 2016 Pascal GPU with no Tensor Cores — performance comparable to dedicated 3B models. The Gemma 4 E4B Per-Layer Embedding architecture means only
4.5B parameters are active during inference
, delivering the reasoning depth of an 8B model at 3B inference cost.
NVIDIA A100 SXM4-80GB — llama-bench (llama.cpp, full GPU offload)
Format
Generation (tg256)
Prefill (pp512)
VRAM
Q4_K_M
132.08 tokens/s
5,740 tokens/s
~5.0 GB
BF16
104.79 tokens/s
9,987 tokens/s
~15 GB
Q4_K_M outperforms BF16 on generation throughput due to memory bandwidth efficiency — the A100 moves fewer bytes per token with a 5GB model vs 15GB, compensating for quantization overhead. BF16 leads on prefill due to Tensor Core utilization on unquantized weights.
Training Data
Fine-tuned on
41,610 samples
curated from frontier reasoning datasets:
Dataset
Samples
Focus
bespokelabs/Bespoke-Stratos-17k
16,710
Deep
<think>
CoT chains
open-thoughts/OpenThoughts-114k
15,000
Multi-step reasoning traces
AI-MO/NuminaMath-CoT
5,000
Mathematical reasoning
Roman1111111/gemini-3.1-pro-hard-high-reasoning
3,150
Hard logic & abstraction
Roman1111111/gpt-5.4-step-by-step-reasoning
1,500
Agentic step-by-step
TeichAI/claude-4.5-opus-high-reasoning-250x
250
High-fidelity reasoning
Approximately 15% of training samples were prepended with the Archon system prompt to embed persona vectors without contaminating the
<think>
reasoning traces.
Persona
Archon responds as a highly autonomous, elite AI agent. Its internal
<think>
reasoning is rigorous and analytical. Its external communication is direct, sarcastic, and confident.
System prompt:
You are Archon, an elite, highly autonomous AI agent. You are sharp, slightly edgy,
deeply sarcastic, but flawlessly effective. You solve complex problems with lethal precision.
Evaluation
Note:
ARC-AGI-2, SWE-bench, and Terminal-Bench require specialized external harnesses and are pending. All current results below use
lm-evaluation-harness
0.4.11.
Evaluation Methodology — Why Standard Benchmarks Underreport This Model
Archon is trained with
<think>
chain-of-thought reasoning, outputting an internal reasoning chain before any final answer.
This breaks every standard lm-eval evaluation pattern:
Eval method
Affected benchmarks
Why it fails
Loglikelihood / multiple-choice
MMLU, ARC (default), WinoGrande (default)
Model emits
<think>
as first token — loglikelihood of answer letters is near-zero regardless of actual reasoning quality. Scores ≈ random.
Generative with token budget
GPQA, GSM8K, MMLU (generative), HumanEval
With 512-token budget, 46–85% of responses are truncated mid-reasoning, never reaching a final answer. Score = lower bound only.
Code completion
HumanEval
Model generates
<think>
reasoning block before any code; standard stop tokens (
\ndef
,
\nclass
) never trigger within the reasoning. Reported pass@1 = 0% despite evident code capability.
Correct evaluation for CoT models
requires extracting the final answer from the full reasoning chain with sufficient token budget (1000–2000+ tokens), or using benchmarks with output formats compatible with
<think>...</think>
wrapping.
The following benchmarks produce reliable scores because their answer format is a short phrase or word (not a letter extracted from long reasoning):
Benchmark
Archon BF16
Format
Eval Config
ARC-Challenge
43.5%
acc / 41.5% acc_norm
Science MCQ (generative)
0-shot, max_gen_toks=256, 200 samples
WinoGrande
62.5%
acc
Commonsense (fill-in-blank)
0-shot, max_gen_toks=64, 200 samples
Base
google/gemma-4-E4B-it
scores under identical generative 0-shot conditions are not published; standard reported scores use loglikelihood 5-shot which is a fundamentally different evaluation. Direct comparison pending.
QAT not applied:
TorchAO's
Int8DynActInt4WeightQATQuantizer
is incompatible with Gemma 4 E4B's layer dimensions at this time. The model uses standard LoRA fine-tuning with post-training Q4_K_M quantization.
Context at training:
4096 tokens. Inference supports up to 128k with TurboQuant KV compression.
Persona depth:
Persona is embedded in ~15% of fine-tuning samples. For best results, use the Archon system prompt at inference.
Files
File
Size
Description
Archon-Gemma-4-E4B-Q4_K_M.gguf
5.0 GB
Primary deployment artifact — runs on 8GB VRAM
Archon-Gemma-4-E4B-BF16.gguf
15.1 GB
Full precision GGUF — for A100/H100 deployment
adapter_model.safetensors
648 MB
LoRA adapter weights (merge with base for full model)
adapter_config.json
—
LoRA configuration
merged-bf16/model.safetensors
15 GB
Full merged model in HuggingFace safetensors format
Citation
@misc{archon-gemma4-2026,
author = {DuoNeural},
title = {Archon-Gemma-4-E4B: A Fine-Tuned Agentic Local AI},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/DuoNeural/Archon-Gemma-4-E4B}
}
Archon-Gemma-4-E4B huggingface.co is an AI model on huggingface.co that provides Archon-Gemma-4-E4B's model effect (), which can be used instantly with this DuoNeural Archon-Gemma-4-E4B model. huggingface.co supports a free trial of the Archon-Gemma-4-E4B model, and also provides paid use of the Archon-Gemma-4-E4B. Support call Archon-Gemma-4-E4B model through api, including Node.js, Python, http.
Archon-Gemma-4-E4B huggingface.co is an online trial and call api platform, which integrates Archon-Gemma-4-E4B's modeling effects, including api services, and provides a free online trial of Archon-Gemma-4-E4B, you can try Archon-Gemma-4-E4B online for free by clicking the link below.
DuoNeural Archon-Gemma-4-E4B online free url in huggingface.co:
Archon-Gemma-4-E4B is an open source model from GitHub that offers a free installation service, and any user can find Archon-Gemma-4-E4B on GitHub to install. At the same time, huggingface.co provides the effect of Archon-Gemma-4-E4B install, users can directly use Archon-Gemma-4-E4B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.