Quark has its own export format and allows FP8 quantized models to be efficiently deployed using the vLLM backend(vLLM-compatible).
Evaluation
Quark currently uses perplexity(PPL) as the evaluation metric for accuracy loss before and after quantization.The specific PPL algorithm can be referenced in the quantize_quark.py.
The quantization evaluation results are conducted in pseudo-quantization mode, which may slightly differ from the actual quantized inference accuracy. These results are provided for reference only.
Evaluation scores
Benchmark
Llama-3.3-70B-Instruct
Llama-3.3-70B-Instruct-FP8-KV(this model)
Perplexity-wikitext2
3.862387180328369
3.9621500968933105
License
Modifications copyright(c) 2024 Advanced Micro Devices,Inc. All rights reserved.
Runs of amd Llama-3.3-70B-Instruct-FP8-KV on huggingface.co
13.8K
Total runs
99
24-hour runs
1.0K
3-day runs
1.3K
7-day runs
1.8K
30-day runs
More Information About Llama-3.3-70B-Instruct-FP8-KV huggingface.co Model
More Llama-3.3-70B-Instruct-FP8-KV license Visit here:
Llama-3.3-70B-Instruct-FP8-KV huggingface.co is an AI model on huggingface.co that provides Llama-3.3-70B-Instruct-FP8-KV's model effect (), which can be used instantly with this amd Llama-3.3-70B-Instruct-FP8-KV model. huggingface.co supports a free trial of the Llama-3.3-70B-Instruct-FP8-KV model, and also provides paid use of the Llama-3.3-70B-Instruct-FP8-KV. Support call Llama-3.3-70B-Instruct-FP8-KV model through api, including Node.js, Python, http.
Llama-3.3-70B-Instruct-FP8-KV huggingface.co is an online trial and call api platform, which integrates Llama-3.3-70B-Instruct-FP8-KV's modeling effects, including api services, and provides a free online trial of Llama-3.3-70B-Instruct-FP8-KV, you can try Llama-3.3-70B-Instruct-FP8-KV online for free by clicking the link below.
amd Llama-3.3-70B-Instruct-FP8-KV online free url in huggingface.co:
Llama-3.3-70B-Instruct-FP8-KV is an open source model from GitHub that offers a free installation service, and any user can find Llama-3.3-70B-Instruct-FP8-KV on GitHub to install. At the same time, huggingface.co provides the effect of Llama-3.3-70B-Instruct-FP8-KV install, users can directly use Llama-3.3-70B-Instruct-FP8-KV installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Llama-3.3-70B-Instruct-FP8-KV install url in huggingface.co: