Bedovyy / Qwen3-32B.w8a8

huggingface.co
Total runs: 32
24-hour runs: -1
7-day runs: -16
30-day runs: -48
Model's Last Updated: May 09 2025
text-generation

Introduction of Qwen3-32B.w8a8

Model Details of Qwen3-32B.w8a8

W8A8 INT8 Quantization of Qwen3-32B

I made this for running on vLLM with Ampere GPU.

On 2xRTX3090, you may set context length to upto 16384 (or 24576 if you use --kv-cache-dtype fp8 ).

Quantization method

Quantized using

## modified based on the code from https://huggingface.co/nytopop/Qwen3-14B.w8a8

from transformers import AutoTokenizer, AutoModelForCausalLM
from datasets import load_dataset
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor.modifiers.smoothquant import SmoothQuantModifier
from llmcompressor.transformers.compression.helpers import calculate_offload_device_map

model_id  = "Qwen/Qwen3-32B"
model_out = "Qwen3-32B.w8a8"

num_samples = 256
max_seq_len = 4096

tokenizer = AutoTokenizer.from_pretrained(model_id)

def preprocess_fn(example):
  return {"text": tokenizer.apply_chat_template(example["messages"], add_generation_prompt=False, tokenize=False)}

ds = load_dataset("neuralmagic/LLM_compression_calibration", split="train")
ds = ds.shuffle().select(range(num_samples))
ds = ds.map(preprocess_fn)

model = AutoModelForCausalLM.from_pretrained(
  model_id,
  device_map="auto",
  torch_dtype="bfloat16",
  max_memory={0: "10GiB", 1:"10GiB", 2:"10GiB", 3:"10GiB", "cpu":"96GiB"},
)

recipe = [
  SmoothQuantModifier(smoothing_strength=0.7),
  GPTQModifier(sequential=True,targets="Linear",scheme="W8A8",ignore=["lm_head"],dampening_frac=0.01),
]

oneshot(
  model=model,
  dataset=ds,
  recipe=recipe,
  max_seq_length=max_seq_len,
  num_calibration_samples=num_samples,
  output_dir=model_out,
)

Runs of Bedovyy Qwen3-32B.w8a8 on huggingface.co

32
Total runs
-1
24-hour runs
0
3-day runs
-16
7-day runs
-48
30-day runs

More Information About Qwen3-32B.w8a8 huggingface.co Model

More Qwen3-32B.w8a8 license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3-32B.w8a8 huggingface.co

Qwen3-32B.w8a8 huggingface.co is an AI model on huggingface.co that provides Qwen3-32B.w8a8's model effect (), which can be used instantly with this Bedovyy Qwen3-32B.w8a8 model. huggingface.co supports a free trial of the Qwen3-32B.w8a8 model, and also provides paid use of the Qwen3-32B.w8a8. Support call Qwen3-32B.w8a8 model through api, including Node.js, Python, http.

Qwen3-32B.w8a8 huggingface.co Url

https://huggingface.co/Bedovyy/Qwen3-32B.w8a8

Bedovyy Qwen3-32B.w8a8 online free

Qwen3-32B.w8a8 huggingface.co is an online trial and call api platform, which integrates Qwen3-32B.w8a8's modeling effects, including api services, and provides a free online trial of Qwen3-32B.w8a8, you can try Qwen3-32B.w8a8 online for free by clicking the link below.

Bedovyy Qwen3-32B.w8a8 online free url in huggingface.co:

https://huggingface.co/Bedovyy/Qwen3-32B.w8a8

Qwen3-32B.w8a8 install

Qwen3-32B.w8a8 is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-32B.w8a8 on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-32B.w8a8 install, users can directly use Qwen3-32B.w8a8 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3-32B.w8a8 install url in huggingface.co:

https://huggingface.co/Bedovyy/Qwen3-32B.w8a8

Url of Qwen3-32B.w8a8

Qwen3-32B.w8a8 huggingface.co Url

Provider of Qwen3-32B.w8a8 huggingface.co

Bedovyy
ORGANIZATIONS

Other API from Bedovyy

huggingface.co

Total runs: 7.8K
Run Growth: 1.5K
Growth Rate: 19.52%
Updated:August 28 2026
huggingface.co

Total runs: 2.0K
Run Growth: -246
Growth Rate: -12.06%
Updated:May 15 2026
huggingface.co

Total runs: 1.6K
Run Growth: 518
Growth Rate: 31.64%
Updated:August 28 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:November 20 2024