EQAQ v2 is the EQC Qwen3.5 4B text-only AWQ target model package used with
SGLang, plus the EAGLE3 draft models used in the local speculative decoding
experiments.
The root model is the target model. The draft directories are EAGLE3 draft
models for SGLang speculative decoding and are not standalone target models.
Expected Performance
These numbers are local measurements from the EQC competition protocol harness,
not an official leaderboard score. The official submission uploaded
successfully, but the evaluation job failed before scoring because the service
could not provision the requested ML compute capacity.
Measured with the EQC latency request shape:
/v1/completions
, logical batch
size 1, 5 warmup runs, 50 measurement runs per category.
The speedup below is computed against a target-only run measured on the same
local machine, not against the fixed baseline constants embedded in the EQC
protocol harness.
Category
Prompt / new tokens
Target-only median
EQAQ v2 median
Local speedup
short
64 / 128
852.58 ms
228.87 ms
3.73x
medium
2048 / 256
1771.02 ms
475.62 ms
3.72x
long
8192 / 256
2179.81 ms
847.43 ms
2.57x
Average local speedup was
3.10x
using the average of category medians
(
1601.14 ms / 517.31 ms
). The older
9.41x
figure comes from dividing by
the EQC harness fixed baseline constants (
2582/5441/6576 ms
) and should not
be interpreted as a speedup over a baseline measured on this machine.
A submission-aligned smoke run with a more conservative single-image setup
measured about
4.39x
against the same fixed protocol constants over 3 runs
per category; it is included only as a packaging/protocol smoke result, not as
the local target-only speedup.
Baseline caveat: the target-only no-spec SGLang server crashed with the default
piecewise CUDA graph path (
NoneType mrope_positions
), so the local
target-only baseline was measured with
--disable-piecewise-cuda-graph
while
keeping the same target model, endpoint, prompt/token protocol, CUDA graph
batch sizes, and core SGLang serving options.
Observed speculative accept rate in the active local SGLang run was low,
roughly
6%
over recent decode batches, so the latency gain should be
understood as the combined effect of SGLang serving settings, CUDA graph, and
speculative decoding rather than high draft acceptance alone.
Local quality
Measured in the same local full protocol run:
Benchmark
Metric
Score
Gate
MMLU-Pro
exact_match, custom-extract
0.6525
0.621
IFEval
inst_level_strict_acc
0.8106
0.814
GPQA-Diamond
exact_match, flexible-extract
0.4293
0.630
The local run passed the latency gate and MMLU-Pro, but did
not
pass the
full quality gate because IFEval was slightly below threshold and GPQA-Diamond
was substantially below threshold. Treat this package as a speed-oriented EQC
artifact, not a confirmed quality-passing competition submission.
More Information About eqaq-v2 huggingface.co Model
eqaq-v2 huggingface.co
eqaq-v2 huggingface.co is an AI model on huggingface.co that provides eqaq-v2's model effect (), which can be used instantly with this NotaMG eqaq-v2 model. huggingface.co supports a free trial of the eqaq-v2 model, and also provides paid use of the eqaq-v2. Support call eqaq-v2 model through api, including Node.js, Python, http.
eqaq-v2 huggingface.co is an online trial and call api platform, which integrates eqaq-v2's modeling effects, including api services, and provides a free online trial of eqaq-v2, you can try eqaq-v2 online for free by clicking the link below.
eqaq-v2 is an open source model from GitHub that offers a free installation service, and any user can find eqaq-v2 on GitHub to install. At the same time, huggingface.co provides the effect of eqaq-v2 install, users can directly use eqaq-v2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.