amd / Llama-3.3-70B-Instruct-FP8-KV

huggingface.co
Total runs: 13.8K
24-hour runs: 99
7-day runs: 1.3K
30-day runs: 1.8K
Model's Last Updated: July 01 2025

Introduction of Llama-3.3-70B-Instruct-FP8-KV

Model Details of Llama-3.3-70B-Instruct-FP8-KV

Llama-3.3-70B-Instruct-FP8-KV

  • Introduction
    This model was built with Llama by applying Quark with calibration samples from Pile dataset.
  • Quantization Stragegy
    • Quantized Layers : All linear layers excluding "lm_head"
    • Weight : FP8 symmetric per-tensor
    • Activation : FP8 symmetric per-tensor
    • KV Cache : FP8 symmetric per-tensor
  • Quick Start
  1. Download and install Quark
  2. Run the quantization script in the example folder using the following command line:
export MODEL_DIR = [local model checkpoint folder] or meta-llama/Llama-3.3-70B-Instruct
python3 quantize_quark.py \
        --model_dir $MODEL_DIR \
        --output_dir $QUANT_MODEL_DIR \
        --quant_scheme w_fp8_a_fp8 \
        --kv_cache_dtype fp8 \
        --num_calib_data 128 \
        --model_export quark_safetensors \
        --no_weight_matrix_merge \
        --multi_gpu \
        --custom_mode fp8
Deployment

Quark has its own export format and allows FP8 quantized models to be efficiently deployed using the vLLM backend(vLLM-compatible).

Evaluation

Quark currently uses perplexity(PPL) as the evaluation metric for accuracy loss before and after quantization.The specific PPL algorithm can be referenced in the quantize_quark.py. The quantization evaluation results are conducted in pseudo-quantization mode, which may slightly differ from the actual quantized inference accuracy. These results are provided for reference only.

Evaluation scores
Benchmark Llama-3.3-70B-Instruct Llama-3.3-70B-Instruct-FP8-KV(this model)
Perplexity-wikitext2 3.862387180328369 3.9621500968933105
License

Modifications copyright(c) 2024 Advanced Micro Devices,Inc. All rights reserved.

Runs of amd Llama-3.3-70B-Instruct-FP8-KV on huggingface.co

13.8K
Total runs
99
24-hour runs
1.0K
3-day runs
1.3K
7-day runs
1.8K
30-day runs

More Information About Llama-3.3-70B-Instruct-FP8-KV huggingface.co Model

More Llama-3.3-70B-Instruct-FP8-KV license Visit here:

https://choosealicense.com/licenses/llama3.3

Llama-3.3-70B-Instruct-FP8-KV huggingface.co

Llama-3.3-70B-Instruct-FP8-KV huggingface.co is an AI model on huggingface.co that provides Llama-3.3-70B-Instruct-FP8-KV's model effect (), which can be used instantly with this amd Llama-3.3-70B-Instruct-FP8-KV model. huggingface.co supports a free trial of the Llama-3.3-70B-Instruct-FP8-KV model, and also provides paid use of the Llama-3.3-70B-Instruct-FP8-KV. Support call Llama-3.3-70B-Instruct-FP8-KV model through api, including Node.js, Python, http.

Llama-3.3-70B-Instruct-FP8-KV huggingface.co Url

https://huggingface.co/amd/Llama-3.3-70B-Instruct-FP8-KV

amd Llama-3.3-70B-Instruct-FP8-KV online free

Llama-3.3-70B-Instruct-FP8-KV huggingface.co is an online trial and call api platform, which integrates Llama-3.3-70B-Instruct-FP8-KV's modeling effects, including api services, and provides a free online trial of Llama-3.3-70B-Instruct-FP8-KV, you can try Llama-3.3-70B-Instruct-FP8-KV online for free by clicking the link below.

amd Llama-3.3-70B-Instruct-FP8-KV online free url in huggingface.co:

https://huggingface.co/amd/Llama-3.3-70B-Instruct-FP8-KV

Llama-3.3-70B-Instruct-FP8-KV install

Llama-3.3-70B-Instruct-FP8-KV is an open source model from GitHub that offers a free installation service, and any user can find Llama-3.3-70B-Instruct-FP8-KV on GitHub to install. At the same time, huggingface.co provides the effect of Llama-3.3-70B-Instruct-FP8-KV install, users can directly use Llama-3.3-70B-Instruct-FP8-KV installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Llama-3.3-70B-Instruct-FP8-KV install url in huggingface.co:

https://huggingface.co/amd/Llama-3.3-70B-Instruct-FP8-KV

Url of Llama-3.3-70B-Instruct-FP8-KV

Llama-3.3-70B-Instruct-FP8-KV huggingface.co Url

Provider of Llama-3.3-70B-Instruct-FP8-KV huggingface.co

amd
ORGANIZATIONS

Other API from amd

huggingface.co

Total runs: 184.5K
Run Growth: 97.3K
Growth Rate: 52.72%
Updated:April 14 2026
huggingface.co

Total runs: 114.2K
Run Growth: -5.4K
Growth Rate: -4.76%
Updated:July 17 2026
huggingface.co

Total runs: 104.0K
Run Growth: 12.6K
Growth Rate: 12.15%
Updated:June 19 2026
huggingface.co

Total runs: 52.6K
Run Growth: -5.8K
Growth Rate: -11.01%
Updated:July 01 2026
huggingface.co

Total runs: 24.6K
Run Growth: 13.9K
Growth Rate: 56.56%
Updated:July 27 2026
huggingface.co

Total runs: 12.4K
Run Growth: 4.6K
Growth Rate: 37.39%
Updated:October 09 2024
huggingface.co

Total runs: 11.2K
Run Growth: 1.2K
Growth Rate: 10.43%
Updated:August 12 2025
huggingface.co

Total runs: 2.5K
Run Growth: -15
Growth Rate: -0.59%
Updated:June 19 2026
huggingface.co

Total runs: 1.3K
Run Growth: -37.3K
Growth Rate: -2875.40%
Updated:June 19 2026
huggingface.co

Total runs: 948
Run Growth: -2.8K
Growth Rate: -295.15%
Updated:May 27 2026
huggingface.co

Total runs: 547
Run Growth: -4.4K
Growth Rate: -812.25%
Updated:June 19 2026
huggingface.co

Total runs: 512
Run Growth: -396
Growth Rate: -77.34%
Updated:November 15 2025