EschaLabs / escha-runtime-qwen3dense

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: September 20 2026

Introduction of escha-runtime-qwen3dense

Model Details of escha-runtime-qwen3dense

Escha Runtime — qwen3dense

By Escha Labs Inc.

The serving runtime for Escha 2-/3-bit ( escha ) quantized models of the qwen3_5 dense architecture (Qwen3.8-27B and siblings). One repo per model architecture, one directory per engine — this architecture currently has one engine, sglang/ .

SGLang — sglang/
Best for everything: single user, teams, agents
Concurrency continuous batching, paged KV, optional radix prefix cache
Tool calls / JSON schema / thinking parser yes
Interface OpenAI-compatible ( /v1/chat/completions , /v1/completions , /v1/models )
Install Python 3.12 venv + CUDA-12 PyTorch, then one wheel

The engine is a fork of SGLang bundled inside the wheel, running the Escha CUDA kernels. No separate sglang install is needed, and none should be present — the wheel ships its own.

Compatible models
Model repo Bits
EschaLabs/Qwen3.8-27B-Escha-W2 2-bit, mixed-rate ( escha )

This runtime targets the qwen3_5 dense architecture. Its wheel also happens to register the eschamoe mixture-of-experts method, so a qwen3_5_moe model will load too — but the tuning, the defaults in sglang/serve.sh and the documentation here are all written for the dense architecture. For a mixture-of-experts model use escha-runtime-qwen3moe , whose defaults are measured on it. A model of a genuinely different architecture will not load — use the matching escha-runtime-<arch> repo.

Quickstart

Full detail, including the per-GPU cookbook and troubleshooting: sglang/INSTALL.md .

python3.12 -m venv .venv && source .venv/bin/activate
pip install -U pip wheel
pip install "torch==2.9.*" --index-url https://download.pytorch.org/whl/cu128   # cu12 torch FIRST
pip install ./sglang/escha-*.whl   # pulls the bundled sglang fork + its full dep closure

hf download EschaLabs/Qwen3.8-27B-Escha-W2 --local-dir ./Qwen3.8-27B-Escha-W2
MODEL=./Qwen3.8-27B-Escha-W2 bash sglang/serve.sh

Then check the stack and the endpoint:

python -c "import torch, escha, sglang; print(torch.cuda.is_available(), hasattr(torch.ops.escha, 'escham_decode_gemv'), escha.__version__)"
curl -s http://127.0.0.1:30000/v1/models | python3 -m json.tool

pip install "torch==2.9.*" is a hard pin, not a suggestion. A bare torch>=2.9 resolves to a newer minor and import escha then fails with undefined symbol: _ZN3c10... — the compiled extension is ABI-linked to libtorch, and that ABI is not stable across PyTorch minors.

Thinking, and why you probably want a budget

This is a reasoning model. With thinking on, the reasoning arrives in reasoning_content and the answer in content — read both , or you will see half the response.

Two per-request levers, both inside chat_template_kwargs (a top-level enable_thinking field is silently ignored):

{ "chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"} }

reasoning_effort is "xhigh" ( the default ), "medium" or "low" ; anything else makes the template raise, which surfaces as an HTTP 400 rather than a silent fallback. It works by injecting one sentence of system instruction — xhigh asks the model to validate assumptions and weigh alternatives, low asks it to keep thinking brief, and medium injects nothing at all , so medium is the neutral, unsteered model rather than a midpoint. It therefore asks for shorter reasoning; it does not bound it. If you are running a benchmark or an agent, set a thinking budget instead, which forces </think> after N reasoning tokens so an answer is always produced: see sglang/INSTALL.md → Bounded thinking and sglang/thinking_budget.py . Without one, the usual failure is finish_reason: "length" with content: null , which a harness scores as wrong rather than as slow .

Requirements
  • NVIDIA GPU, compute capability 8.0–12.0 (Ampere → Blackwell), Linux x86-64 with glibc ≥ 2.28 . The kernel launch route auto-selects per GPU; you never set it. Per-architecture and per-VRAM launch recipes: sglang/INSTALL.md → Running on your GPU .
  • Python 3.12 (the wheel is cp312 -only) + CUDA-12 PyTorch 2.9.x . The wheel handles every other dependency.
  • A working host C compiler and Python dev headers. Triton JIT-compiles a small shim at CUDA-graph capture time — this is separate from ptxas and from a CUDA toolkit, so "driver only" does not cover it. On slim container images a stripped libisl breaks cc1 while gcc --version still succeeds, and the failure surfaces ~40 s in as a gcc CalledProcessError inside cuda_graph_runner.py — which reads like a runtime bug and is not. Preflight in sglang/INSTALL.md .
  • 24 GB VRAM for the shipped defaults (65,536-token context, ~8–9 concurrent streams) with a ~10.15 GB model. Note the stream ceiling: MAMBA_RATIO=0.3 sizes the recurrent-state pool, which clamps max_running_requests to 8–9 on a 24 GB card, so the 12 / 16 entries in the default CUDA_GRAPH_BS are dropped and never captured . To serve more streams raise MAXREQ / MAXMAMBA with MEM — the throughput recipe is in the model card . 16 GB should fit at a reduced context — the cookbook has a recipe, but we have not run it.
Changelog

1.2.0 (2026-08-21) — tensor parallelism ( --tp-size N ) now works. The escha parameter class pins its own weight loader, which meant sglang's TP slicing never ran and every rank kept the whole checkpoint (rank 0 died with weight must have shape (dim, width) ). It now slices per rank, including the fused-on-disk GDN in_proj_qkv , which is split into its three sub-projections first.

Single-GPU users are unaffected. Every new code path is gated on world_size > 1 ; at --tp-size 1 the loader is byte-for-byte what 1.1.1 did. Verified as an identical shard layout and byte-identical greedy output.

TP > 1 is new and lightly tested — treat it as experimental. It was contributed and validated by @ginerJuanUdesa on 2× RTX 3090 (symmetric 6.02 GB/rank, coherent greedy output). We have one GPU and could not reproduce it , and no numerical equivalence check against --tp-size 1 has been run yet. If you use it for evaluation, sanity-check a benchmark against the single-GPU numbers first. Note that a multi-rank all-reduce reorders float accumulation, so TP > 1 output is not expected to match TP = 1 bit-for-bit even when correct.

On Ampere/Ada/Hopper you can add DETERMINISTIC=1 to remove that reduction-order variance if you want a stricter comparison.

1.1.1 (2026-08-21) — process_weights_after_loading now takes the rank's device instead of a hardcoded cuda:0 . The hardcode put every 2-bit buffer on cuda:0 while the input tensor sat on the server's actual device, so any run not on device 0 — --tp-size > 1 , or a single-GPU launch with --base-gpu-id N and no CUDA_VISIBLE_DEVICES — hit an illegal memory access on the first forward , behind a traceback that pointed at the kernel rather than at the cause. Bit-identical wherever cuda:0 was already correct, which is every configuration serve.sh ships. Reported with a diagnosis and a fix by @ginerJuanUdesa .

1.1.0 (2026-08-20) — first wheel with the dense ( escha ) serving path; the 1.0.x wheels registered eschamoe only, so a dense checkpoint failed at registry lookup.

Known limitations
  • Serving throughput has been measured on three cards — RTX 5090 (32 GB, sm_120), RTX 4090 (24 GB, sm_89) and RTX 3090 (24 GB, sm_86); the 5090 and 3090 by independent evaluators working only from these docs. The 4090 is also where the model's thinking-on benchmarks (GPQA-Diamond, LiveCodeBench) were produced; its thinking-off commonsense suite ran on an L40. The 16 GB and 40 GB+ tiers in the cookbook remain configuration guidance derived from the model size and the wheel's architecture coverage, not measurements . Per-GPU numbers live on the model card .
  • On Ampere the auto-selected kernel route is the slower one at batch 1. ESCHA_ROUTE resolves to lovelace on sm_80/sm_86, but forcing ESCHA_ROUTE=blackwell measured 1.72× faster single-stream on an RTX 3090 (23.6 → 40.7 tok/s, TPOT 42.4 → 24.6 ms) with identical output. The two routes are bit-identical launch geometries, so this is safe to set; the gain is batch-1-only (parity at 2–16). Serving one user on Ampere? Set it.
  • DETERMINISTIC=1 fails on consumer Blackwell (sm_120). The deterministic attention kernel requests 104 KB of shared memory per block, above the sm_120 limit, and the server exits during startup. It works on Ampere, Ada and Hopper.
  • Greedy output is not bit-reproducible across requests. Batch composition changes fp16 accumulation order, so a near-tie can flip and a long reasoning chain diverges from there. Two identical greedy requests may return different text. Use DETERMINISTIC=1 when you need reproducibility, and never A/B two configurations by diffing one generation.
  • CUDA-graph batch sizes are capped at 32 , the decode kernel's maximum M ( torch.ops.escha.escham_decode_gemv_max_m() ). The shipped default list stops at 16 because that is where aggregate throughput peaks on a 4090; capture at 24 / 32 works and is worth it if you serve that many streams. Past 32 a batch falls through to a large-M path meant for prefill, so the runtime refuses to capture it rather than bake in the wrong kernel.
  • ATTN_BACKEND=triton is required on consumer Blackwell (RTX 50-series). The default flashinfer backend asserts on this hybrid architecture at sm_120. The assertion names three acceptable backends — triton , trtllm_mha , fa4 — of which only triton has been run on this model. Note that sm_120 shows steeper long-prompt decode decay than sm_89 (88.5% vs 96.3% of short-prompt rate at a 5,000-token prompt); the attention path is the obvious suspect and nobody has run the A/B that would confirm it.
  • The model's stock chat template raises on a few malformed conversations (a missing user query, a system message in the wrong position). That surfaces as an HTTP 400 from the template, not as a server fault.
License

Everything here is released under the Apache License, Version 2.0 — see LICENSE . All bundled third-party code is permissive (Apache-2.0 / MIT / BSD-3-Clause) — no copyleft . Full texts and the component inventory: THIRD_PARTY_LICENSES/ . Model weights are not in this repo and carry their own license in the model repository.

Runs of EschaLabs escha-runtime-qwen3dense on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About escha-runtime-qwen3dense huggingface.co Model

More escha-runtime-qwen3dense license Visit here:

https://choosealicense.com/licenses/apache-2.0

escha-runtime-qwen3dense huggingface.co

escha-runtime-qwen3dense huggingface.co is an AI model on huggingface.co that provides escha-runtime-qwen3dense's model effect (), which can be used instantly with this EschaLabs escha-runtime-qwen3dense model. huggingface.co supports a free trial of the escha-runtime-qwen3dense model, and also provides paid use of the escha-runtime-qwen3dense. Support call escha-runtime-qwen3dense model through api, including Node.js, Python, http.

escha-runtime-qwen3dense huggingface.co Url

https://huggingface.co/EschaLabs/escha-runtime-qwen3dense

EschaLabs escha-runtime-qwen3dense online free

escha-runtime-qwen3dense huggingface.co is an online trial and call api platform, which integrates escha-runtime-qwen3dense's modeling effects, including api services, and provides a free online trial of escha-runtime-qwen3dense, you can try escha-runtime-qwen3dense online for free by clicking the link below.

EschaLabs escha-runtime-qwen3dense online free url in huggingface.co:

https://huggingface.co/EschaLabs/escha-runtime-qwen3dense

escha-runtime-qwen3dense install

escha-runtime-qwen3dense is an open source model from GitHub that offers a free installation service, and any user can find escha-runtime-qwen3dense on GitHub to install. At the same time, huggingface.co provides the effect of escha-runtime-qwen3dense install, users can directly use escha-runtime-qwen3dense installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

escha-runtime-qwen3dense install url in huggingface.co:

https://huggingface.co/EschaLabs/escha-runtime-qwen3dense

Url of escha-runtime-qwen3dense

escha-runtime-qwen3dense huggingface.co Url

Provider of escha-runtime-qwen3dense huggingface.co

EschaLabs
ORGANIZATIONS

Other API from EschaLabs