This repository contains an offline-FP8 version of
llmfan46/Qwen3.5-9B-ultra-heretic
, prepared for serving on a single NVIDIA
L40S
with SGLang.
This is not a GGUF export and not an online quantization recipe. It is a saved Hugging Face-style quantized checkpoint produced offline with
llmcompressor
, then validated on Modal with SGLang.
What this repo contains
model.safetensors
config.json
generation_config.json
tokenizer files
multimodal processor files
quantization_manifest.json
recipe.yaml
The artifact was saved in a form that SGLang can load directly from disk.
The original Heretic model card and the upstream Qwen model card remain the authoritative references for training provenance and behavior of the unquantized model.
Quantization notes
Quantization method: offline FP8
Primary tool:
llmcompressor
Scheme:
FP8_DYNAMIC
Target hardware during quantization:
1x NVIDIA L40S
Target runtime during validation: SGLang on
1x NVIDIA L40S
The quantized artifact includes the processor and tokenizer sidecars required for multimodal loading.
Compatibility notes
Two compatibility fixes were necessary to make the saved checkpoint load cleanly in SGLang:
video_preprocessor_config.json
was added so the multimodal processor stack could initialize correctly.
tokenizer_config.json
was normalized to use
Qwen2TokenizerFast
instead of
TokenizersBackend
.
Those fixes are already included in this repo.
Measured serving results on Modal
These numbers were measured with SGLang on a single
L40S
, loading this saved checkpoint directly from a Modal Volume.
Long-context prefill benchmark:
Context window:
262,144
Request size used for testing: about
90%
of context
Prompt tokens sent:
235,939
Uncached 5-run batch:
Average prompt TPS:
5932.34
Median prompt TPS:
6070.03
Warmed steady-state average excluding run 1:
6120.02
Average peak VRAM:
40.160 GiB
Max peak VRAM:
40.161 GiB
Important interpretation:
These are single-request prefill measurements with
max_tokens=1
.
They are not decode throughput measurements.
They are not multi-user throughput measurements.
Suggested SGLang launch
The recommended launch pattern is a normal load of the saved checkpoint:
Qwen3.5-9B-ultra-heretic-fp8 huggingface.co is an AI model on huggingface.co that provides Qwen3.5-9B-ultra-heretic-fp8's model effect (), which can be used instantly with this apothic Qwen3.5-9B-ultra-heretic-fp8 model. huggingface.co supports a free trial of the Qwen3.5-9B-ultra-heretic-fp8 model, and also provides paid use of the Qwen3.5-9B-ultra-heretic-fp8. Support call Qwen3.5-9B-ultra-heretic-fp8 model through api, including Node.js, Python, http.
Qwen3.5-9B-ultra-heretic-fp8 huggingface.co is an online trial and call api platform, which integrates Qwen3.5-9B-ultra-heretic-fp8's modeling effects, including api services, and provides a free online trial of Qwen3.5-9B-ultra-heretic-fp8, you can try Qwen3.5-9B-ultra-heretic-fp8 online for free by clicking the link below.
apothic Qwen3.5-9B-ultra-heretic-fp8 online free url in huggingface.co:
Qwen3.5-9B-ultra-heretic-fp8 is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.5-9B-ultra-heretic-fp8 on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.5-9B-ultra-heretic-fp8 install, users can directly use Qwen3.5-9B-ultra-heretic-fp8 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3.5-9B-ultra-heretic-fp8 install url in huggingface.co: