This is a quantized version of
gpt-oss-20b-BF16
created by AMD using LLM Compressor (compressed-tensors) for ZenDNN-optimized CPU inference.
Quantization
The model was quantized from
gpt-oss-20b-BF16
using LLM Compressor via the Round-to-Nearest (RTN) algorithm. This reduces the model weights from 40 GiB to 21 GiB on disk (~47% reduction).
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.modeling.gpt_oss import convert_model_for_quantization_gptoss
model_id = "unsloth/gpt-oss-20b-BF16"
output_dir = "./gpt-oss-20b-BF16-w8a8"# Step 1: Load the BF16 model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="cpu",
trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
# Step 2: Expand fused gpt-oss MoE experts into individual nn.Linear modules# so QuantizationModifier quantizes every expert.
convert_model_for_quantization_gptoss(model)
# Step 3: Define the W8A8 recipe
recipe = QuantizationModifier(
scheme="W8A8",
targets=["Linear"],
ignore=[
"lm_head",
r"re:.*\.router$",
r"re:.*\.router\..*",
r"re:.*\.gate$",
r"re:.*\.mlp\.gate$",
],
)
# Step 4: One-shot quantize and save in compressed-tensors format
oneshot(
model=model,
recipe=recipe,
tokenizer=tokenizer,
output_dir=output_dir,
trust_remote_code_model=True,
)
# Smoke test
inputs = tokenizer("What are we having for dinner?", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=30)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Quick Start
Use with vLLM
from vllm import LLM, SamplingParams
model = LLM(
model="amd/gpt-oss-20b-BF16-w8a8-llmcompressor-v0.10.0.2",
dtype="bfloat16",
)
sampling_params = SamplingParams(temperature=0.7, max_tokens=256)
outputs = model.generate(["Hello, how are you?"], sampling_params)
print(outputs[0].outputs[0].text)
For optimal performance, set
LD_PRELOAD
with
libomp.so
(LLVM OpenMP) or
libiomp5.so
(Intel OpenMP):
# Using LLVM OpenMP (llvmopenmp)export LD_PRELOAD=$(find /path/to/env -name "libomp.so" | head -1)
# Or using Intel OpenMP (libiomp)export LD_PRELOAD=$(find /path/to/env -name "libiomp5.so" | head -1)
Note:
Set
LD_PRELOAD
before launching vLLM or any inference script.
Evaluation
The model was evaluated against the BF16 (unquantized) baseline on standard benchmarks using
lm-evaluation-harness
with the vLLM engine.
gpt-oss-20b-BF16-w8a8-llmcompressor huggingface.co is an AI model on huggingface.co that provides gpt-oss-20b-BF16-w8a8-llmcompressor's model effect (), which can be used instantly with this amd gpt-oss-20b-BF16-w8a8-llmcompressor model. huggingface.co supports a free trial of the gpt-oss-20b-BF16-w8a8-llmcompressor model, and also provides paid use of the gpt-oss-20b-BF16-w8a8-llmcompressor. Support call gpt-oss-20b-BF16-w8a8-llmcompressor model through api, including Node.js, Python, http.
gpt-oss-20b-BF16-w8a8-llmcompressor huggingface.co is an online trial and call api platform, which integrates gpt-oss-20b-BF16-w8a8-llmcompressor's modeling effects, including api services, and provides a free online trial of gpt-oss-20b-BF16-w8a8-llmcompressor, you can try gpt-oss-20b-BF16-w8a8-llmcompressor online for free by clicking the link below.
amd gpt-oss-20b-BF16-w8a8-llmcompressor online free url in huggingface.co:
gpt-oss-20b-BF16-w8a8-llmcompressor is an open source model from GitHub that offers a free installation service, and any user can find gpt-oss-20b-BF16-w8a8-llmcompressor on GitHub to install. At the same time, huggingface.co provides the effect of gpt-oss-20b-BF16-w8a8-llmcompressor install, users can directly use gpt-oss-20b-BF16-w8a8-llmcompressor installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
gpt-oss-20b-BF16-w8a8-llmcompressor install url in huggingface.co: