Introduction of Llama-3.3-70B-Instruct-quantized.w4a16
Model Details of Llama-3.3-70B-Instruct-quantized.w4a16
Llama-3.3-70B-Instruct-quantized.w4a16
Model Overview
Model Architecture:
Meta-Llama-3.1
Input:
Text
Output:
Text
Model Optimizations:
Weight quantization:
INT4
Intended Use Cases:
Intended for commercial and research use in multiple languages. Instruction tuned text only models are intended for assistant-like chat, whereas pretrained models can be adapted for a variety of natural language generation tasks. The Llama 3.3 model also supports the ability to leverage the outputs of its models to improve other models including synthetic data generation and distillation. The Llama 3.3 Community License allows for these use cases.
Out-of-scope:
Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in any other way that is prohibited by the Acceptable Use Policy and Llama 3.3 Community License. Use in languages beyond English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.
Release Date:
12/11/2024
Version:
1.0
License(s):
llama3.3
Model Developers:
Red Hat (Neural Magic)
Model Optimizations
This model was obtained by quantizing the weights of
Llama-3.3-70B-Instruct
to INT4 data type.
This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%.
Only the weights of the linear operators within transformers blocks are quantized.
Weights are quantized using a symmetric per-group scheme, with group size 128.
The
GPTQ
algorithm is applied for quantization, as implemented in the
llm-compressor
library.
Deployment
This model can be deployed efficiently using the
vLLM
backend, as shown in the example below.
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer
model_id = "RedHatAI/Llama-3.3-70B-Instruct-quantized.w4a16"
number_gpus = 1
sampling_params = SamplingParams(temperature=0.6, top_p=0.9, max_tokens=256)
tokenizer = AutoTokenizer.from_pretrained(model_id)
messages = [
{"role": "system", "content": "You are a pirate chatbot who always responds in pirate speak!"},
{"role": "user", "content": "Who are you?"},
]
prompts = tokenizer.apply_chat_template(messages, tokenize=False)
llm = LLM(model=model_id, tensor_parallel_size=number_gpus)
outputs = llm.generate(prompts, sampling_params)
generated_text = outputs[0].outputs[0].text
print(generated_text)
vLLM aslo supports OpenAI-compatible serving. See the
documentation
for more details.
# Download model from Red Hat Registry via docker# Note: This downloads the model to ~/.cache/instructlab/models unless --model-dir is specified.
ilab model download --repository docker://registry.redhat.io/rhelai1/llama-3-3-70b-instruct-quantized-w4a16:1.5
# Serve model via ilab
ilab model serve --model-path ~/.cache/instructlab/models/llama-3-3-70b-instruct-quantized-w4a16
# Chat with model
ilab model chat --model ~/.cache/instructlab/models/llama-3-3-70b-instruct-quantized-w4a16
# Setting up vllm server with ServingRuntime# Save as: vllm-servingruntime.yaml
apiVersion: serving.kserve.io/v1alpha1
kind: ServingRuntime
metadata:
name: vllm-cuda-runtime # OPTIONAL CHANGE: set a unique name
annotations:
openshift.io/display-name: vLLM NVIDIA GPU ServingRuntime for KServe
opendatahub.io/recommended-accelerators: '["nvidia.com/gpu"]'
labels:
opendatahub.io/dashboard: 'true'
spec:
annotations:
prometheus.io/port: '8080'
prometheus.io/path: '/metrics'
multiModel: false
supportedModelFormats:
- autoSelect: true
name: vLLM
containers:
- name: kserve-container
image: quay.io/modh/vllm:rhoai-2.20-cuda # CHANGE if needed. If AMD: quay.io/modh/vllm:rhoai-2.20-rocm
command:
- python
- -m
- vllm.entrypoints.openai.api_server
args:
- "--port=8080"
- "--model=/mnt/models"
- "--served-model-name={{.Name}}"
env:
- name: HF_HOME
value: /tmp/hf_home
ports:
- containerPort: 8080
protocol: TCP
# Attach model to vllm server. This is an NVIDIA template# Save as: inferenceservice.yaml
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
annotations:
openshift.io/display-name: llama-3-3-70b-instruct-quantized-w4a16 # OPTIONAL CHANGE
serving.kserve.io/deploymentMode: RawDeployment
name: llama-3-3-70b-instruct-quantized-w4a16 # specify model name. This value will be used to invoke the model in the payload
labels:
opendatahub.io/dashboard: 'true'
spec:
predictor:
maxReplicas: 1
minReplicas: 1
model:
modelFormat:
name: vLLM
name: ''
resources:
limits:
cpu: '2'# this is model specific
memory: 8Gi # this is model specific
nvidia.com/gpu: '1'# this is accelerator specific
requests: # same comment for this block
cpu: '1'
memory: 4Gi
nvidia.com/gpu: '1'
runtime: vllm-cuda-runtime # must match the ServingRuntime name above
storageUri: oci://registry.redhat.io/rhelai1/modelcar-llama-3-3-70b-instruct-quantized-w4a16:1.5
tolerations:
- effect: NoSchedule
key: nvidia.com/gpu
operator: Exists
# make sure first to be in the project where you want to deploy the model# oc project <project-name># apply both resources to run model# Apply the ServingRuntime
oc apply -f vllm-servingruntime.yaml
# Apply the InferenceService
oc apply -f qwen-inferenceservice.yaml
# Replace <inference-service-name> and <cluster-ingress-domain> below:# - Run `oc get inferenceservice` to find your URL if unsure.# Call the server using curl:
curl https://<inference-service-name>-predictor-default.<domain>/v1/chat/completions
-H "Content-Type: application/json" \
-d '{ "model": "llama-3-3-70b-instruct-quantized-w4a16", "stream": true, "stream_options": { "include_usage": true }, "max_tokens": 1, "messages": [ { "role": "user", "content": "How can a bee fly when its wings are so small?" } ]}'
Creation details
This model was created with [llm-compressor](https://github.com/vllm-project/llm-compressor) by running the code snippet below.
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor.transformers import oneshot
from datasets import load_dataset
# Load model
model_stub = "meta-llama/Llama-3.3-70B-Instruct"
model_name = model_stub.split("/")[-1]
num_samples = 1024
max_seq_len = 8192
tokenizer = AutoTokenizer.from_pretrained(model_stub)
model = AutoModelForCausalLM.from_pretrained(
model_stub,
device_map="auto",
torch_dtype="auto",
)
defpreprocess_fn(example):
return {"text": tokenizer.apply_chat_template(example["messages"], add_generation_prompt=False, tokenize=False)}
ds = load_dataset("neuralmagic/LLM_compression_calibration", split="train")
ds = ds.map(preprocess_fn)
# Configure the quantization algorithm and scheme
recipe = GPTQModifier(
targets="Linear",
scheme="W4A16",
ignore=["lm_head"],
sequential_targets=["LlamaDecoderLayer"],
dampening_frac=0.01,
)
# Apply quantization
oneshot(
model=model,
dataset=ds,
recipe=recipe,
max_seq_length=max_seq_len,
num_calibration_samples=num_samples,
)
# Save to disk in compressed-tensors format
save_path = model_name + "-quantized.w4a16"
model.save_pretrained(save_path)
tokenizer.save_pretrained(save_path)
print(f"Model and tokenizer saved to: {save_path}")
Evaluation
This model was evaluated on the well-known OpenLLM v1, HumanEval, and HumanEval+ benchmarks.
In all cases, model outputs were generated with the
vLLM
engine.
Llama-3.3-70B-Instruct-quantized.w4a16 huggingface.co is an AI model on huggingface.co that provides Llama-3.3-70B-Instruct-quantized.w4a16's model effect (), which can be used instantly with this RedHatAI Llama-3.3-70B-Instruct-quantized.w4a16 model. huggingface.co supports a free trial of the Llama-3.3-70B-Instruct-quantized.w4a16 model, and also provides paid use of the Llama-3.3-70B-Instruct-quantized.w4a16. Support call Llama-3.3-70B-Instruct-quantized.w4a16 model through api, including Node.js, Python, http.
Llama-3.3-70B-Instruct-quantized.w4a16 huggingface.co is an online trial and call api platform, which integrates Llama-3.3-70B-Instruct-quantized.w4a16's modeling effects, including api services, and provides a free online trial of Llama-3.3-70B-Instruct-quantized.w4a16, you can try Llama-3.3-70B-Instruct-quantized.w4a16 online for free by clicking the link below.
RedHatAI Llama-3.3-70B-Instruct-quantized.w4a16 online free url in huggingface.co:
Llama-3.3-70B-Instruct-quantized.w4a16 is an open source model from GitHub that offers a free installation service, and any user can find Llama-3.3-70B-Instruct-quantized.w4a16 on GitHub to install. At the same time, huggingface.co provides the effect of Llama-3.3-70B-Instruct-quantized.w4a16 install, users can directly use Llama-3.3-70B-Instruct-quantized.w4a16 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Llama-3.3-70B-Instruct-quantized.w4a16 install url in huggingface.co: