The Latest AIs, every day
AIs with the most favorites on Toolify
AIs with the highest website traffic (monthly visits)
AI Tools by Apps
Discover the Discord of AI
AI Tools by browser extensions
GPTs from GPT Store
Discover The Best Model For AI
Top AI lists by month and monthly visits.
Top AI lists by category and monthly visits.
Top AI lists by region and monthly visits.
Top AI lists by source and monthly visits.
Top AI lists by revenue and real traffic.

qwen3dense
The serving runtime for
Escha
2-/3-bit (
escha
) quantized models of the
qwen3_5
dense
architecture
(Qwen3.8-27B and siblings). One repo per model architecture, one directory per
engine — this architecture currently has
one
engine,
sglang/
.
SGLang
—
sglang/
|
|
|---|---|
| Best for | everything: single user, teams, agents |
| Concurrency | continuous batching, paged KV, optional radix prefix cache |
| Tool calls / JSON schema / thinking parser | yes |
| Interface |
OpenAI-compatible (
/v1/chat/completions
,
/v1/completions
,
/v1/models
)
|
| Install | Python 3.12 venv + CUDA-12 PyTorch, then one wheel |
The engine is a fork of
SGLang
bundled inside the wheel,
running the Escha CUDA kernels. No separate
sglang
install is needed, and none should be
present — the wheel ships its own.
| Model repo | Bits |
|---|---|
| EschaLabs/Qwen3.8-27B-Escha-W2 |
2-bit, mixed-rate (
escha
)
|
This runtime targets the
qwen3_5dense architecture. Its wheel also happens to register theeschamoemixture-of-experts method, so aqwen3_5_moemodel will load too — but the tuning, the defaults insglang/serve.shand the documentation here are all written for the dense architecture. For a mixture-of-experts model useescha-runtime-qwen3moe, whose defaults are measured on it. A model of a genuinely different architecture will not load — use the matchingescha-runtime-<arch>repo.
Full detail, including the per-GPU cookbook and troubleshooting:
sglang/INSTALL.md
.
python3.12 -m venv .venv && source .venv/bin/activate
pip install -U pip wheel
pip install "torch==2.9.*" --index-url https://download.pytorch.org/whl/cu128 # cu12 torch FIRST
pip install ./sglang/escha-*.whl # pulls the bundled sglang fork + its full dep closure
hf download EschaLabs/Qwen3.8-27B-Escha-W2 --local-dir ./Qwen3.8-27B-Escha-W2
MODEL=./Qwen3.8-27B-Escha-W2 bash sglang/serve.sh
Then check the stack and the endpoint:
python -c "import torch, escha, sglang; print(torch.cuda.is_available(), hasattr(torch.ops.escha, 'escham_decode_gemv'), escha.__version__)"
curl -s http://127.0.0.1:30000/v1/models | python3 -m json.tool
pip install "torch==2.9.*"is a hard pin, not a suggestion. A baretorch>=2.9resolves to a newer minor andimport eschathen fails withundefined symbol: _ZN3c10...— the compiled extension is ABI-linked to libtorch, and that ABI is not stable across PyTorch minors.
This is a reasoning model. With thinking on, the reasoning arrives in
reasoning_content
and the
answer in
content
—
read both
, or you will see half the response.
Two per-request levers, both inside
chat_template_kwargs
(a
top-level
enable_thinking
field
is silently ignored):
{ "chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"} }
reasoning_effort
is
"xhigh"
(
the default
),
"medium"
or
"low"
; anything else makes the
template raise, which surfaces as an HTTP 400 rather than a silent fallback. It works by injecting
one sentence of system instruction —
xhigh
asks the model to validate assumptions and weigh
alternatives,
low
asks it to keep thinking brief, and
medium
injects nothing at all
, so
medium
is the neutral, unsteered model rather than a midpoint. It therefore
asks
for shorter
reasoning; it does not bound it. If you are running a benchmark
or an agent, set a
thinking budget
instead, which forces
</think>
after N reasoning tokens so
an answer is always produced: see
sglang/INSTALL.md
→ Bounded thinking
and
sglang/thinking_budget.py
. Without one, the usual failure is
finish_reason: "length"
with
content: null
, which a harness scores as
wrong
rather than as
slow
.
sglang/INSTALL.md
→ Running on your GPU
.
cp312
-only) +
CUDA-12 PyTorch 2.9.x
. The wheel handles every
other dependency.
ptxas
and from a CUDA toolkit, so "driver only"
does not cover it. On slim container images a stripped
libisl
breaks
cc1
while
gcc --version
still succeeds, and the failure surfaces ~40 s in as a
gcc
CalledProcessError
inside
cuda_graph_runner.py
— which reads like a runtime bug and is not.
Preflight in
sglang/INSTALL.md
.
MAMBA_RATIO=0.3
sizes the recurrent-state pool, which
clamps
max_running_requests
to 8–9 on a 24 GB card, so the
12
/
16
entries in the default
CUDA_GRAPH_BS
are
dropped and never captured
. To serve more streams raise
MAXREQ
/
MAXMAMBA
with
MEM
— the throughput recipe is in the
model card
.
16 GB should fit at a reduced context — the cookbook has a recipe, but we have not run it.
1.2.0
(2026-08-21) —
tensor parallelism (
--tp-size N
) now works.
The escha
parameter class pins its own weight loader, which meant sglang's TP slicing never ran and
every rank kept the whole checkpoint (rank 0 died with
weight must have shape (dim, width)
). It now slices per rank, including the fused-on-disk GDN
in_proj_qkv
,
which is split into its three sub-projections first.
Single-GPU users are unaffected. Every new code path is gated on
world_size > 1; at--tp-size 1the loader is byte-for-byte what 1.1.1 did. Verified as an identical shard layout and byte-identical greedy output.TP > 1 is new and lightly tested — treat it as experimental. It was contributed and validated by @ginerJuanUdesa on 2× RTX 3090 (symmetric 6.02 GB/rank, coherent greedy output). We have one GPU and could not reproduce it , and no numerical equivalence check against
--tp-size 1has been run yet. If you use it for evaluation, sanity-check a benchmark against the single-GPU numbers first. Note that a multi-rank all-reduce reorders float accumulation, so TP > 1 output is not expected to match TP = 1 bit-for-bit even when correct.On Ampere/Ada/Hopper you can add
DETERMINISTIC=1to remove that reduction-order variance if you want a stricter comparison.
1.1.1
(2026-08-21) —
process_weights_after_loading
now takes the rank's device
instead of a hardcoded
cuda:0
. The hardcode put every 2-bit buffer on
cuda:0
while
the input tensor sat on the server's actual device, so
any run not on device 0 —
--tp-size > 1
, or a single-GPU launch with
--base-gpu-id N
and no
CUDA_VISIBLE_DEVICES
— hit an illegal memory access on the first forward
, behind a
traceback that pointed at the kernel rather than at the cause. Bit-identical wherever
cuda:0
was already correct, which is every configuration
serve.sh
ships. Reported
with a diagnosis and a fix by
@ginerJuanUdesa
.
1.1.0
(2026-08-20) — first wheel with the dense (
escha
) serving path; the 1.0.x
wheels registered
eschamoe
only, so a dense checkpoint failed at registry lookup.
ESCHA_ROUTE
resolves to
lovelace
on sm_80/sm_86, but forcing
ESCHA_ROUTE=blackwell
measured
1.72×
faster single-stream on an RTX 3090
(23.6 → 40.7 tok/s, TPOT 42.4 → 24.6 ms) with identical
output. The two routes are bit-identical launch geometries, so this is safe to set; the gain is
batch-1-only (parity at 2–16). Serving one user on Ampere? Set it.
DETERMINISTIC=1
fails on consumer Blackwell (sm_120).
The deterministic attention kernel
requests 104 KB of shared memory per block, above the sm_120 limit, and the server exits during
startup. It works on Ampere, Ada and Hopper.
DETERMINISTIC=1
when you need
reproducibility, and never A/B two configurations by diffing one generation.
torch.ops.escha.escham_decode_gemv_max_m()
). The shipped default list stops at 16 because
that is where aggregate throughput peaks on a 4090; capture at
24
/
32
works and is worth it if
you serve that many streams. Past 32 a batch falls through to a large-M path meant for prefill,
so the runtime refuses to capture it rather than bake in the wrong kernel.
ATTN_BACKEND=triton
is required on consumer Blackwell (RTX 50-series).
The default
flashinfer backend asserts on this hybrid architecture at sm_120. The assertion names three
acceptable backends —
triton
,
trtllm_mha
,
fa4
— of which only
triton
has been run on
this model. Note that sm_120 shows steeper long-prompt decode decay than sm_89 (88.5% vs 96.3%
of short-prompt rate at a 5,000-token prompt); the attention path is the obvious suspect and
nobody has run the A/B that would confirm it.
Everything here is released under the
Apache License, Version 2.0
— see
LICENSE
.
All bundled third-party code is permissive (Apache-2.0 / MIT / BSD-3-Clause) —
no copyleft
.
Full texts and the component inventory:
THIRD_PARTY_LICENSES/
. Model weights are
not
in this repo and carry
their own license in the model repository.
escha-runtime-qwen3dense huggingface.co is an AI model on huggingface.co that provides escha-runtime-qwen3dense's model effect (), which can be used instantly with this EschaLabs escha-runtime-qwen3dense model. huggingface.co supports a free trial of the escha-runtime-qwen3dense model, and also provides paid use of the escha-runtime-qwen3dense. Support call escha-runtime-qwen3dense model through api, including Node.js, Python, http.
escha-runtime-qwen3dense huggingface.co is an online trial and call api platform, which integrates escha-runtime-qwen3dense's modeling effects, including api services, and provides a free online trial of escha-runtime-qwen3dense, you can try escha-runtime-qwen3dense online for free by clicking the link below.
escha-runtime-qwen3dense is an open source model from GitHub that offers a free installation service, and any user can find escha-runtime-qwen3dense on GitHub to install. At the same time, huggingface.co provides the effect of escha-runtime-qwen3dense install, users can directly use escha-runtime-qwen3dense installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
