import torch
from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.gptq import GPTQModifier
model_id = "microsoft/Phi-4-reasoning-plus"
output_dir = "./Phi-4-reasoning-plus-w4a16-llmcompressor-v0.12.0"
NUM_CALIBRATION_SAMPLES = 128
MAX_SEQUENCE_LENGTH = 2048# Step 1: Load the BF16 model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map="cpu",
trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
# Step 2: Load calibration data. GPTQ is data-driven: it needs real activations# to build the per-layer Hessians used to compensate the rounding error.
ds = load_dataset(
"HuggingFaceH4/ultrachat_200k",
split=f"train_sft[:{NUM_CALIBRATION_SAMPLES}]",
)
ds = ds.map(
lambda example: {"text": "\n".join(m["content"] for m in example["messages"])},
remove_columns=ds.column_names,
)
# Step 3: Define the W4A16 GPTQ recipe
recipe = GPTQModifier(scheme="W4A16", targets="Linear", ignore=["lm_head"])
# Step 4: One-shot quantize with calibration and save in compressed-tensors format
oneshot(
model=model,
dataset=ds,
recipe=recipe,
max_seq_length=MAX_SEQUENCE_LENGTH,
tokenizer=tokenizer,
output_dir=output_dir,
trust_remote_code_model=True,
)
# Smoke test
inputs = tokenizer("What are we having for dinner?", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=30)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Quick Start
Use with vLLM
from vllm import LLM, SamplingParams
model = LLM(
model="amd/Phi-4-reasoning-plus-w4a16-llmcompressor-v0.12.0",
dtype="bfloat16",
)
sampling_params = SamplingParams(temperature=0.7, max_tokens=256)
outputs = model.generate(["Hello, how are you?"], sampling_params)
print(outputs[0].outputs[0].text)
For optimal performance, set
LD_PRELOAD
with
libomp.so
(LLVM OpenMP) or
libiomp5.so
(Intel OpenMP):
# Using LLVM OpenMP (llvmopenmp)export LD_PRELOAD=$(find /path/to/env -name "libomp.so" | head -1)
# Or using Intel OpenMP (libiomp)export LD_PRELOAD=$(find /path/to/env -name "libiomp5.so" | head -1)
Note:
Set
LD_PRELOAD
before launching vLLM or any inference script.
Evaluation
The model was evaluated against the BF16 (unquantized) baseline on standard benchmarks using
lm-evaluation-harness
with the vLLM engine.
Phi-4-reasoning-plus-w4a16-llmcompressor huggingface.co is an AI model on huggingface.co that provides Phi-4-reasoning-plus-w4a16-llmcompressor's model effect (), which can be used instantly with this amd Phi-4-reasoning-plus-w4a16-llmcompressor model. huggingface.co supports a free trial of the Phi-4-reasoning-plus-w4a16-llmcompressor model, and also provides paid use of the Phi-4-reasoning-plus-w4a16-llmcompressor. Support call Phi-4-reasoning-plus-w4a16-llmcompressor model through api, including Node.js, Python, http.
Phi-4-reasoning-plus-w4a16-llmcompressor huggingface.co is an online trial and call api platform, which integrates Phi-4-reasoning-plus-w4a16-llmcompressor's modeling effects, including api services, and provides a free online trial of Phi-4-reasoning-plus-w4a16-llmcompressor, you can try Phi-4-reasoning-plus-w4a16-llmcompressor online for free by clicking the link below.
amd Phi-4-reasoning-plus-w4a16-llmcompressor online free url in huggingface.co:
Phi-4-reasoning-plus-w4a16-llmcompressor is an open source model from GitHub that offers a free installation service, and any user can find Phi-4-reasoning-plus-w4a16-llmcompressor on GitHub to install. At the same time, huggingface.co provides the effect of Phi-4-reasoning-plus-w4a16-llmcompressor install, users can directly use Phi-4-reasoning-plus-w4a16-llmcompressor installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Phi-4-reasoning-plus-w4a16-llmcompressor install url in huggingface.co: