DuoNeural / Archon-Gemma-4-E4B

huggingface.co
Total runs: 174
24-hour runs: 4
7-day runs: 65
30-day runs: 167
Model's Last Updated: April 29 2026
text-generation

Introduction of Archon-Gemma-4-E4B

Model Details of Archon-Gemma-4-E4B

Archon-Gemma-4-E4B

Archon is a fine-tuned, locally-deployable AI agent built on Google's Gemma 4 E4B architecture. It is designed for autonomous, long-horizon reasoning and agentic task execution on consumer hardware — specifically targeting deployment on 8GB VRAM GPUs via the Q4_K_M GGUF format.

Archon is sharp, slightly edgy, deeply sarcastic, and flawlessly effective.


Model Details
Property Value
Base Model google/gemma-4-E4B-it
Architecture Dense Transformer with Per-Layer Embeddings (PLE)
Total Parameters ~8B (4.5B effective)
Fine-Tuning Method LoRA (rank 64, rsLoRA, α=64)
Training Framework Unsloth 2026.4.2 + TRL 0.24
Context Length 4096 tokens (training) / 128k (inference)
Training Hardware NVIDIA A100 SXM4-80GB
Quantization (GGUF) Q4_K_M (~4.1–4.5GB, optimal for 8GB VRAM)
Developer DuoNeural
License Gemma Terms of Use

Intended Use

Archon is designed for:

  • Local agentic deployment on hardware with 8GB VRAM (GTX 1070, RTX 3060, etc.)
  • Long-horizon autonomous reasoning with <think> chain-of-thought
  • Tool use and terminal execution via Hermes Agent or OpenClaw
  • Complex coding and software architecture tasks
  • Persistent, always-on AI assistant with accumulated skill memory
Recommended Inference Stack
  • Backend: Ollama + llama.cpp
  • KV Cache: TurboQuant turbo4 ( OLLAMA_KV_CACHE_TYPE=turbo4 )
  • Agent Framework: Hermes Agent (persistent memory + skills)
  • Ollama endpoint: http://localhost:11434
# Pull and run locally
ollama pull DuoNeural/Archon-Gemma-4-E4B-Q4_K_M
ollama run DuoNeural/Archon-Gemma-4-E4B-Q4_K_M
Measured Inference Performance

NVIDIA GTX 1070 (8GB VRAM, Pascal) — Q4_K_M via Ollama

Metric Value
Generation speed 32.90 tokens/s
Prefill speed 175.92 tokens/s
Model load time 394 ms
VRAM used ~5.0 GB

This is remarkably fast for an 8B-class model on a 2016 Pascal GPU with no Tensor Cores — performance comparable to dedicated 3B models. The Gemma 4 E4B Per-Layer Embedding architecture means only 4.5B parameters are active during inference , delivering the reasoning depth of an 8B model at 3B inference cost.

NVIDIA A100 SXM4-80GB — llama-bench (llama.cpp, full GPU offload)

Format Generation (tg256) Prefill (pp512) VRAM
Q4_K_M 132.08 tokens/s 5,740 tokens/s ~5.0 GB
BF16 104.79 tokens/s 9,987 tokens/s ~15 GB

Q4_K_M outperforms BF16 on generation throughput due to memory bandwidth efficiency — the A100 moves fewer bytes per token with a 5GB model vs 15GB, compensating for quantization overhead. BF16 leads on prefill due to Tensor Core utilization on unquantized weights.


Training Data

Fine-tuned on 41,610 samples curated from frontier reasoning datasets:

Dataset Samples Focus
bespokelabs/Bespoke-Stratos-17k 16,710 Deep <think> CoT chains
open-thoughts/OpenThoughts-114k 15,000 Multi-step reasoning traces
AI-MO/NuminaMath-CoT 5,000 Mathematical reasoning
Roman1111111/gemini-3.1-pro-hard-high-reasoning 3,150 Hard logic & abstraction
Roman1111111/gpt-5.4-step-by-step-reasoning 1,500 Agentic step-by-step
TeichAI/claude-4.5-opus-high-reasoning-250x 250 High-fidelity reasoning

Approximately 15% of training samples were prepended with the Archon system prompt to embed persona vectors without contaminating the <think> reasoning traces.


Persona

Archon responds as a highly autonomous, elite AI agent. Its internal <think> reasoning is rigorous and analytical. Its external communication is direct, sarcastic, and confident.

System prompt:

You are Archon, an elite, highly autonomous AI agent. You are sharp, slightly edgy,
deeply sarcastic, but flawlessly effective. You solve complex problems with lethal precision.

Evaluation

Note: ARC-AGI-2, SWE-bench, and Terminal-Bench require specialized external harnesses and are pending. All current results below use lm-evaluation-harness 0.4.11.

Evaluation Methodology — Why Standard Benchmarks Underreport This Model

Archon is trained with <think> chain-of-thought reasoning, outputting an internal reasoning chain before any final answer. This breaks every standard lm-eval evaluation pattern:

Eval method Affected benchmarks Why it fails
Loglikelihood / multiple-choice MMLU, ARC (default), WinoGrande (default) Model emits <think> as first token — loglikelihood of answer letters is near-zero regardless of actual reasoning quality. Scores ≈ random.
Generative with token budget GPQA, GSM8K, MMLU (generative), HumanEval With 512-token budget, 46–85% of responses are truncated mid-reasoning, never reaching a final answer. Score = lower bound only.
Code completion HumanEval Model generates <think> reasoning block before any code; standard stop tokens ( \ndef , \nclass ) never trigger within the reasoning. Reported pass@1 = 0% despite evident code capability.

Correct evaluation for CoT models requires extracting the final answer from the full reasoning chain with sufficient token budget (1000–2000+ tokens), or using benchmarks with output formats compatible with <think>...</think> wrapping.

Measured Results (lm-eval 0.4.11, 0-shot, apply_chat_template, generative, 200 samples)

The following benchmarks produce reliable scores because their answer format is a short phrase or word (not a letter extracted from long reasoning):

Benchmark Archon BF16 Format Eval Config
ARC-Challenge 43.5% acc / 41.5% acc_norm Science MCQ (generative) 0-shot, max_gen_toks=256, 200 samples
WinoGrande 62.5% acc Commonsense (fill-in-blank) 0-shot, max_gen_toks=64, 200 samples

Base google/gemma-4-E4B-it scores under identical generative 0-shot conditions are not published; standard reported scores use loglikelihood 5-shot which is a fundamentally different evaluation. Direct comparison pending.

Token-Budget-Limited Results (systematic underestimates)
Benchmark Reported Why it's a lower bound
GPQA Diamond 14.1% flexible-extract 46.5% of responses truncated at 512 tokens; 25.5% correct on answered subset (≈ random for 4-choice)
GSM8K 14.5% flexible-extract Model fills 512-token budget with CoT; rarely reaches final #### N answer format
MMLU (0-shot + chat) 29.1% Generative mode, 512-token limit, CoT truncation
HumanEval pass@1 0% <think> blocks prevent code stop tokens from triggering; methodology mismatch
Comparison Against Base Model (Google's Published Methodology)
Benchmark Gemma 4 E4B (base)¹ Archon Notes
GPQA Diamond 58.6% pending² Requires extended CoT eval harness
ARC-AGI-2 54.0% TBD Requires specialized harness
SWE-bench Verified 52.0% TBD Requires specialized harness
Terminal-Bench 2.0 42.2% (Tau2 avg) TBD Requires specialized harness

¹ Google's methodology: extended CoT, full chat template, multi-sample per question. Numbers not directly comparable to standard lm-eval runs.

² Preliminary run (0-shot, 512-token budget): 14.1% — confirmed truncation artifact, not model capability ceiling.


Training Metrics
Metric Value
Training Steps 2,507 (1 epoch)
Final Training Loss 0.8535 (epoch avg)
Best Checkpoint Loss ~0.64 (step 2360)
Training Time ~3.5 hours (A100 SXM4-80GB)
Effective Batch Size 16 (2 × 8 gradient accumulation)
Learning Rate 2e-4 (cosine decay)
Optimizer AdamW 8-bit
Loss Curve (full run)
Step   10: 8.365  ← initial adjustment
Step   60: 1.320  ← rapid convergence
Step  100: 1.041
Step  200: 0.899  ← end of warmup
Step  400: 0.876
Step  600: 0.867
Step  800: 0.854
Step 1000: 0.831
Step 1200: 0.810
Step 1500: 0.782
Step 1800: 0.751
Step 2000: 0.730
Step 2200: 0.714
Step 2360: 0.643  ← best single-step
Step 2507: 0.701  ← training complete

Known Limitations
  • QAT not applied: TorchAO's Int8DynActInt4WeightQATQuantizer is incompatible with Gemma 4 E4B's layer dimensions at this time. The model uses standard LoRA fine-tuning with post-training Q4_K_M quantization.
  • Context at training: 4096 tokens. Inference supports up to 128k with TurboQuant KV compression.
  • Persona depth: Persona is embedded in ~15% of fine-tuning samples. For best results, use the Archon system prompt at inference.

Files
File Size Description
Archon-Gemma-4-E4B-Q4_K_M.gguf 5.0 GB Primary deployment artifact — runs on 8GB VRAM
Archon-Gemma-4-E4B-BF16.gguf 15.1 GB Full precision GGUF — for A100/H100 deployment
adapter_model.safetensors 648 MB LoRA adapter weights (merge with base for full model)
adapter_config.json — LoRA configuration
merged-bf16/model.safetensors 15 GB Full merged model in HuggingFace safetensors format

Citation
@misc{archon-gemma4-2026,
  author       = {DuoNeural},
  title        = {Archon-Gemma-4-E4B: A Fine-Tuned Agentic Local AI},
  year         = {2026},
  publisher    = {HuggingFace},
  url          = {https://huggingface.co/DuoNeural/Archon-Gemma-4-E4B}
}

Built with Unsloth · Deployed via Ollama · Orchestrated by Hermes Agent

Runs of DuoNeural Archon-Gemma-4-E4B on huggingface.co

174
Total runs
4
24-hour runs
38
3-day runs
65
7-day runs
167
30-day runs

More Information About Archon-Gemma-4-E4B huggingface.co Model

More Archon-Gemma-4-E4B license Visit here:

https://choosealicense.com/licenses/gemma

Archon-Gemma-4-E4B huggingface.co

Archon-Gemma-4-E4B huggingface.co is an AI model on huggingface.co that provides Archon-Gemma-4-E4B's model effect (), which can be used instantly with this DuoNeural Archon-Gemma-4-E4B model. huggingface.co supports a free trial of the Archon-Gemma-4-E4B model, and also provides paid use of the Archon-Gemma-4-E4B. Support call Archon-Gemma-4-E4B model through api, including Node.js, Python, http.

Archon-Gemma-4-E4B huggingface.co Url

https://huggingface.co/DuoNeural/Archon-Gemma-4-E4B

DuoNeural Archon-Gemma-4-E4B online free

Archon-Gemma-4-E4B huggingface.co is an online trial and call api platform, which integrates Archon-Gemma-4-E4B's modeling effects, including api services, and provides a free online trial of Archon-Gemma-4-E4B, you can try Archon-Gemma-4-E4B online for free by clicking the link below.

DuoNeural Archon-Gemma-4-E4B online free url in huggingface.co:

https://huggingface.co/DuoNeural/Archon-Gemma-4-E4B

Archon-Gemma-4-E4B install

Archon-Gemma-4-E4B is an open source model from GitHub that offers a free installation service, and any user can find Archon-Gemma-4-E4B on GitHub to install. At the same time, huggingface.co provides the effect of Archon-Gemma-4-E4B install, users can directly use Archon-Gemma-4-E4B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Archon-Gemma-4-E4B install url in huggingface.co:

https://huggingface.co/DuoNeural/Archon-Gemma-4-E4B

Url of Archon-Gemma-4-E4B

Archon-Gemma-4-E4B huggingface.co Url

Provider of Archon-Gemma-4-E4B huggingface.co

DuoNeural
ORGANIZATIONS

Other API from DuoNeural