cognitivecomputations / DeepSeek-V3-0324-AWQ

huggingface.co
Total runs: 4.1K
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: March 29 2025
text-generation

Introduction of DeepSeek-V3-0324-AWQ

Model Details of DeepSeek-V3-0324-AWQ

DeepSeek V3 0324 AWQ

AWQ of DeepSeek V3 0324.

Quantized by Eric Hartford and v2ray .

This quant modified some of the model code to fix an overflow issue when using float16.

To serve using vLLM with 8x 80GB GPUs, use the following command:

VLLM_USE_V1=0 VLLM_WORKER_MULTIPROC_METHOD=spawn VLLM_MARLIN_USE_ATOMIC_ADD=1 python -m vllm.entrypoints.openai.api_server --host 0.0.0.0 --port 12345 --max-model-len 65536 --max-seq-len-to-capture 65536 --enable-chunked-prefill --enable-prefix-caching --trust-remote-code --tensor-parallel-size 8 --gpu-memory-utilization 0.95 --served-model-name deepseek-chat --model cognitivecomputations/DeepSeek-V3-0324-AWQ

You can download the wheel I built for PyTorch 2.6, Python 3.12 by clicking here , the benchmark below was done with this wheel, it contains 2 PR merges which boosted performance a lot.

TPS Per Request
GPU \ Batch Input Output B: 1 I: 2 O: 2K B: 32 I: 4K O: 256 B: 1 I: 63K O: 2K Prefill
8x H100/H200 61.5 30.1 54.3 4732.2
4x H200 58.4 19.8 53.7 2653.1
8x A100 80GB 45.5 11.9 7.3 2435.5

Note:

  • The A100 config is extremely slow on high context is caused by FlashMLA not supporting anything below Hopper GPUs (H200, H100, H800, H20), before it's supported, vLLM will use the Triton implementation which is extremely slow on high context. Thus it's best to serve this model with either 8x H100 or 4x H200.
  • All 3 types of GPU are SXM form factor.
  • Inference speed will be better than FP8 at low batch size but worse than FP8 at high batch size, this is the nature of low bit quantization.
  • vLLM supports MLA for AWQ now, you can run this model with full context length on just 8x 80GB GPUs.

Runs of cognitivecomputations DeepSeek-V3-0324-AWQ on huggingface.co

4.1K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About DeepSeek-V3-0324-AWQ huggingface.co Model

More DeepSeek-V3-0324-AWQ license Visit here:

https://choosealicense.com/licenses/mit

DeepSeek-V3-0324-AWQ huggingface.co

DeepSeek-V3-0324-AWQ huggingface.co is an AI model on huggingface.co that provides DeepSeek-V3-0324-AWQ's model effect (), which can be used instantly with this cognitivecomputations DeepSeek-V3-0324-AWQ model. huggingface.co supports a free trial of the DeepSeek-V3-0324-AWQ model, and also provides paid use of the DeepSeek-V3-0324-AWQ. Support call DeepSeek-V3-0324-AWQ model through api, including Node.js, Python, http.

cognitivecomputations DeepSeek-V3-0324-AWQ online free

DeepSeek-V3-0324-AWQ huggingface.co is an online trial and call api platform, which integrates DeepSeek-V3-0324-AWQ's modeling effects, including api services, and provides a free online trial of DeepSeek-V3-0324-AWQ, you can try DeepSeek-V3-0324-AWQ online for free by clicking the link below.

cognitivecomputations DeepSeek-V3-0324-AWQ online free url in huggingface.co:

https://huggingface.co/cognitivecomputations/DeepSeek-V3-0324-AWQ

DeepSeek-V3-0324-AWQ install

DeepSeek-V3-0324-AWQ is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V3-0324-AWQ on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V3-0324-AWQ install, users can directly use DeepSeek-V3-0324-AWQ installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

DeepSeek-V3-0324-AWQ install url in huggingface.co:

https://huggingface.co/cognitivecomputations/DeepSeek-V3-0324-AWQ

Url of DeepSeek-V3-0324-AWQ

Provider of DeepSeek-V3-0324-AWQ huggingface.co

cognitivecomputations
ORGANIZATIONS

Other API from cognitivecomputations