lainlives / OmniCoder-9B-bnb-4bit

huggingface.co
Total runs: 5
24-hour runs: 0
7-day runs: 0
30-day runs: -16
Model's Last Updated: March 19 2026
text-generation

Introduction of OmniCoder-9B-bnb-4bit

Model Details of OmniCoder-9B-bnb-4bit

Tesslate/OmniCoder-9B (Quantized)

Description

This model is a quantized version of the original model Tesslate/OmniCoder-9B .

Quantization Details
  • Quantization Type : int4
  • bnb_4bit_quant_type : nf4
  • bnb_4bit_use_double_quant : True
  • bnb_4bit_compute_dtype : bfloat16
  • bnb_4bit_quant_storage : int8

📄 Original Model Information

OmniCoder

OmniCoder-9B

A 9B coding agent fine-tuned on 425K agentic trajectories.

License Base Model GGUF

!! 3/12/26 Update -> Install For Your Coding Agents

Get Started | Benchmarks | GGUF Downloads


Overview

OmniCoder-9B is a 9-billion parameter coding agent model built by Tesslate , fine-tuned on top of Qwen3.5-9B 's hybrid architecture (Gated Delta Networks interleaved with standard attention). It was trained on 425,000+ curated agentic coding trajectories spanning real-world software engineering tasks, tool use, terminal operations, and multi-step reasoning.

The training data was specifically built from Claude Opus 4.6 agentic and coding reasoning traces , targeting scaffolding patterns from Claude Code, OpenCode, Codex, and Droid. The dataset includes successful trajectories from models like Claude Opus 4.6, GPT-5.4, GPT-5.3-Codex, and Gemini 3.1 Pro.

The model shows strong agentic behavior: it recovers from errors (read-before-write), responds to LSP diagnostics, and uses proper edit diffs instead of full rewrites. These patterns were learned directly from the real-world agent trajectories it was trained on.

Key Features
  • Trained on Frontier Agent Traces : Built from Claude Opus 4.6, GPT-5.3-Codex, GPT-5.4, and Gemini 3.1 Pro agentic coding trajectories across Claude Code, OpenCode, Codex, and Droid scaffolding
  • Hybrid Architecture : Inherits Qwen3.5's Gated Delta Networks interleaved with standard attention for efficient long-context processing
  • 262K Native Context : Full 262,144 token context window, extensible to 1M+
  • Error Recovery : Learns read-before-write patterns, responds to LSP diagnostics, and applies minimal edit diffs instead of full rewrites
  • Thinking Mode : Supports <think>...</think> reasoning chains for complex problem decomposition
  • Apache 2.0 : Fully open weights, no restrictions

Benchmarks
Benchmark OmniCoder-9B Qwen3.5-9B Qwen3-Next-80B GPT-OSS-120B GPT-OSS-20B GLM-4.7-Flash GLM 4.7 Claude Haiku 4.5
AIME 2025 (pass@5) 90 91.7 91.6
GPQA Diamond (pass@1) 83.8 81.7 77.2 80.1 71.5 73
GPQA Diamond (pass@3) 86.4
Terminal-Bench 2.0 23.6 14.6 33.4 27
  • GPQA Diamond pass@1: 83.8% (166/198). +2.1 points over the Qwen3.5-9B base model (81.7). At pass@3: 86.4 (171/198).
  • AIME 2025 pass@5: 90% (27/30).
  • Terminal-Bench 2.0: 23.6% (21/89). +8.99 points (+61% improvement) over the Qwen3.5-9B base model (14.6%, 13/89).

Quickstart
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Tesslate/OmniCoder-9B"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

messages = [
    {"role": "system", "content": "You are a helpful coding assistant."},
    {"role": "user", "content": "Write a Python function to find the longest common subsequence of two strings."},
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=2048, temperature=0.6, top_p=0.95, top_k=20)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True))
vLLM
vllm serve Tesslate/OmniCoder-9B --tensor-parallel-size 1 --max-model-len 65536
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="token")
response = client.chat.completions.create(
    model="Tesslate/OmniCoder-9B",
    messages=[{"role": "user", "content": "Explain the difference between a mutex and a semaphore."}],
    temperature=0.6,
)
print(response.choices[0].message.content)
llama.cpp (GGUF)
llama-cli --hf-repo Tesslate/OmniCoder-9B-GGUF --hf-file omnicoder-9b-q4_k_m.gguf -p "Your prompt" -c 8192

All quantizations: Tesslate/OmniCoder-9B-GGUF


Training Details
Base Model Qwen3.5-9B
Method LoRA SFT (r=64, alpha=32)
Dataset 425K agentic trajectories from 5 sources
Packing Sample packing with 99.35% efficiency
Hardware 4x NVIDIA H200 (DDP)
Framework Axolotl
Precision bf16
Optimizer AdamW (lr=2e-4, cosine schedule)

Architecture

OmniCoder inherits Qwen3.5-9B's hybrid architecture:

  • Gated Delta Networks : Linear attention layers interleaved with standard attention for efficient long-range dependencies
  • VLM Backbone : Built on Qwen3_5ForConditionalGeneration

Recommended Sampling Parameters
Parameter Value
Temperature 0.6
Top-P 0.95
Top-K 20
Presence Penalty 0.0

For agentic / tool-calling tasks, consider lower temperature (0.2-0.4) for more deterministic behavior.


Limitations
  • Performance on non-English tasks has not been extensively evaluated
  • Tool-calling format is flexible but works best with the scaffolding patterns seen in training

Acknowledgments

Special thanks to the Axolotl team and the discussion in axolotl#3453 for helping get Qwen3.5 packing support working.


Citation
@misc{omnicoder2025,
  title={OmniCoder-9B: A Frontier Open Coding Agent},
  author={Tesslate},
  year={2025},
  url={https://huggingface.co/Tesslate/OmniCoder-9B}
}

Built by Tesslate

Runs of lainlives OmniCoder-9B-bnb-4bit on huggingface.co

5
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
-16
30-day runs

More Information About OmniCoder-9B-bnb-4bit huggingface.co Model

More OmniCoder-9B-bnb-4bit license Visit here:

https://choosealicense.com/licenses/apache-2.0

OmniCoder-9B-bnb-4bit huggingface.co

OmniCoder-9B-bnb-4bit huggingface.co is an AI model on huggingface.co that provides OmniCoder-9B-bnb-4bit's model effect (), which can be used instantly with this lainlives OmniCoder-9B-bnb-4bit model. huggingface.co supports a free trial of the OmniCoder-9B-bnb-4bit model, and also provides paid use of the OmniCoder-9B-bnb-4bit. Support call OmniCoder-9B-bnb-4bit model through api, including Node.js, Python, http.

OmniCoder-9B-bnb-4bit huggingface.co Url

https://huggingface.co/lainlives/OmniCoder-9B-bnb-4bit

lainlives OmniCoder-9B-bnb-4bit online free

OmniCoder-9B-bnb-4bit huggingface.co is an online trial and call api platform, which integrates OmniCoder-9B-bnb-4bit's modeling effects, including api services, and provides a free online trial of OmniCoder-9B-bnb-4bit, you can try OmniCoder-9B-bnb-4bit online for free by clicking the link below.

lainlives OmniCoder-9B-bnb-4bit online free url in huggingface.co:

https://huggingface.co/lainlives/OmniCoder-9B-bnb-4bit

OmniCoder-9B-bnb-4bit install

OmniCoder-9B-bnb-4bit is an open source model from GitHub that offers a free installation service, and any user can find OmniCoder-9B-bnb-4bit on GitHub to install. At the same time, huggingface.co provides the effect of OmniCoder-9B-bnb-4bit install, users can directly use OmniCoder-9B-bnb-4bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

OmniCoder-9B-bnb-4bit install url in huggingface.co:

https://huggingface.co/lainlives/OmniCoder-9B-bnb-4bit

Url of OmniCoder-9B-bnb-4bit

OmniCoder-9B-bnb-4bit huggingface.co Url

Provider of OmniCoder-9B-bnb-4bit huggingface.co

lainlives
ORGANIZATIONS

Other API from lainlives

huggingface.co

Total runs: 442
Run Growth: 302
Growth Rate: 68.33%
Updated:May 08 2026
huggingface.co

Total runs: 101
Run Growth: 0
Growth Rate: 0.00%
Updated:May 12 2026
huggingface.co

Total runs: 65
Run Growth: -172
Growth Rate: -264.62%
Updated:March 15 2026
huggingface.co

Total runs: 62
Run Growth: 0
Growth Rate: 0.00%
Updated:May 12 2026
huggingface.co

Total runs: 18
Run Growth: 18
Growth Rate: 100.00%
Updated:May 11 2026
huggingface.co

Total runs: 17
Run Growth: 3
Growth Rate: 17.65%
Updated:May 10 2026
huggingface.co

Total runs: 12
Run Growth: 12
Growth Rate: 100.00%
Updated:May 08 2026
huggingface.co

Total runs: 11
Run Growth: 11
Growth Rate: 100.00%
Updated:May 07 2026
huggingface.co

Total runs: 10
Run Growth: -2
Growth Rate: -20.00%
Updated:February 02 2026
huggingface.co

Total runs: 7
Run Growth: 5
Growth Rate: 71.43%
Updated:May 08 2026
huggingface.co

Total runs: 5
Run Growth: 3
Growth Rate: 60.00%
Updated:April 08 2026
huggingface.co

Total runs: 4
Run Growth: 2
Growth Rate: 50.00%
Updated:July 30 2026
huggingface.co

Total runs: 4
Run Growth: 4
Growth Rate: 100.00%
Updated:May 08 2026
huggingface.co

Total runs: 3
Run Growth: 0
Growth Rate: 0.00%
Updated:March 05 2026
huggingface.co

Total runs: 3
Run Growth: -3
Growth Rate: -100.00%
Updated:May 12 2026
huggingface.co

Total runs: 2
Run Growth: 1
Growth Rate: 50.00%
Updated:March 05 2026
huggingface.co

Total runs: 2
Run Growth: 0
Growth Rate: 0.00%
Updated:March 05 2026
huggingface.co

Total runs: 2
Run Growth: 0
Growth Rate: 0.00%
Updated:March 05 2026
huggingface.co

Total runs: 2
Run Growth: -1
Growth Rate: -50.00%
Updated:May 07 2026