skt / A.X-K2-NVFP4

huggingface.co
Total runs: 194.4K
24-hour runs: 3.4K
7-day runs: 27.7K
30-day runs: 192.7K
Model's Last Updated: August 07 2026
text-generation

Introduction of A.X-K2-NVFP4

Model Details of A.X-K2-NVFP4

A.X K2

A.X Logo

🤗 Models | 🖥️ Github | 📄 Technical Report

Model Summary

A.X K2 is a large-scale Mixture-of-Experts (MoE) language model trained from scratch as a high-performance, agentic foundation model, and the successor to A.X K1. The model contains 688 billion total parameters , with 33 billion active parameters , delivering strong reasoning and instruction-following performance while maintaining practical inference efficiency.

Through a Think-Fusion training recipe, a single unified model supports both a thinking mode for complex problem solving and a non-thinking mode for concise, low-latency responses, allowing the user to trade quality for cost on a per-request basis.

A.X K2 is developed as part of the Korean government's Sovereign AI foundation model project, aiming to build a frontier-scale model with deep understanding of the Korean language and culture.

A.X K2 NVFP4

A.X K2 NVFP4 is an NVFP4-quantized version of A.X K2 , a 688B-parameter Mixture-of-Experts language model developed by SK Telecom.

This checkpoint applies NVFP4 W4A4 quantization to the routed experts , while keeping the remaining modules in FP8 or BF16. It retains performance comparable to the FP8 checkpoint while reducing the model memory footprint by approximately half.

The model can be served on a single node with 4 NVIDIA B200 GPUs using the A.X K2–enabled SKT-AI vLLM fork.

Quantization
Module Precision
Routed experts NVFP4 W4A4
Attention, router, shared expert, and dense layer FP8
Embedding and LM head BF16
Usage

Install the A.X K2–enabled vLLM fork:

git clone -b axk2-v0.23.0 https://github.com/SKT-AI/vllm.git
cd vllm
pip install -e .

Enable the FlashInfer NVFP4 MoE kernels:

export VLLM_USE_FLASHINFER_MOE_FP4=1
export FLASHINFER_DISABLE_VERSION_CHECK=1

Serve the model:

vllm serve skt/A.X-K2-NVFP4 \
  --served-model-name A.X-K2-NVFP4 \
  --host 0.0.0.0 \
  --port 8000 \
  --tensor-parallel-size 4 \
  --enable-expert-parallel \
  --attention-backend FLASHINFER_MLA_SPARSE \
  --reasoning-parser deepseek_v3

The A.X K2 implementation is registered directly in the SKT-AI vLLM branch, so --trust-remote-code is not required.

To enable automatic tool calling, add:

--enable-auto-tool-choice \
--tool-call-parser hermes
Thinking Mode

Thinking mode can be controlled per request through chat_template_kwargs .

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "A.X-K2-NVFP4",
    "messages": [
      {
        "role": "user",
        "content": "Explain how speculative decoding works."
      }
    ],
    "max_tokens": 1024,
    "chat_template_kwargs": {
      "enable_thinking": true
    }
  }'

Set enable_thinking to false for a direct non-thinking response.

Long Context

The checkpoint includes the YaRN configuration required for context lengths of up to 256K tokens . No separate RoPE override is required when using the provided config.json .

The serving context can be reduced with --max-model-len when longer inputs are not required.

Needle-in-a-Haystack (NIAH)

NIAH probes exact fact retrieval by inserting a target fact (the "needle") at varying depths within a long context and asking the model to recover it. Under zero-shot YaRN scaling (scaling factors of 1, 2, and 4 for 128K, 256K, and 512K, respectively), A.X K2 attains a perfect retrieval score at every context length and needle depth—and does so even after NVFP4 (experts-only W4A4) quantization. The heatmaps below show the 256K and 512K results.

A.X K2 NIAH fact retrieval at 256K context under NVFP4 (experts-only W4A4) A.X K2 NIAH fact retrieval at 512K context under NVFP4 (experts-only W4A4)

NIAH fact retrieval across context length (x-axis) and needle depth (y-axis): 256K (YaRN factor 2, top) and 512K (YaRN factor 4, bottom), under NVFP4 (experts-only W4A4) quantization. A.X K2 scores a perfect 100 at every position.

License

A.X K2 is released under the Apache License 2.0.

Citation
@techreport{axk2-2026,
  title       = {A.X K2 Technical Report},
  author      = {SK Telecom},
  year        = {2026},
  institution = {SK Telecom},
  url         = {https://github.com/SKT-AI/A.X-K2/blob/main/A_X_K2_Tech_Report.pdf}
}

Runs of skt A.X-K2-NVFP4 on huggingface.co

194.4K
Total runs
3.4K
24-hour runs
17.8K
3-day runs
27.7K
7-day runs
192.7K
30-day runs

More Information About A.X-K2-NVFP4 huggingface.co Model

More A.X-K2-NVFP4 license Visit here:

https://choosealicense.com/licenses/apache-2.0

A.X-K2-NVFP4 huggingface.co

A.X-K2-NVFP4 huggingface.co is an AI model on huggingface.co that provides A.X-K2-NVFP4's model effect (), which can be used instantly with this skt A.X-K2-NVFP4 model. huggingface.co supports a free trial of the A.X-K2-NVFP4 model, and also provides paid use of the A.X-K2-NVFP4. Support call A.X-K2-NVFP4 model through api, including Node.js, Python, http.

A.X-K2-NVFP4 huggingface.co Url

https://huggingface.co/skt/A.X-K2-NVFP4

skt A.X-K2-NVFP4 online free

A.X-K2-NVFP4 huggingface.co is an online trial and call api platform, which integrates A.X-K2-NVFP4's modeling effects, including api services, and provides a free online trial of A.X-K2-NVFP4, you can try A.X-K2-NVFP4 online for free by clicking the link below.

skt A.X-K2-NVFP4 online free url in huggingface.co:

https://huggingface.co/skt/A.X-K2-NVFP4

A.X-K2-NVFP4 install

A.X-K2-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find A.X-K2-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of A.X-K2-NVFP4 install, users can directly use A.X-K2-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

A.X-K2-NVFP4 install url in huggingface.co:

https://huggingface.co/skt/A.X-K2-NVFP4

Url of A.X-K2-NVFP4

A.X-K2-NVFP4 huggingface.co Url

Provider of A.X-K2-NVFP4 huggingface.co

skt
ORGANIZATIONS

Other API from skt

huggingface.co

Total runs: 312.0K
Run Growth: -117.1K
Growth Rate: -37.52%
Updated:September 24 2021
huggingface.co

Total runs: 46.5K
Run Growth: -52.6K
Growth Rate: -113.26%
Updated:September 01 2026
huggingface.co

Total runs: 13.9K
Run Growth: 62
Growth Rate: 0.45%
Updated:July 23 2025
huggingface.co

Total runs: 11.9K
Run Growth: -350
Growth Rate: -2.95%
Updated:June 12 2025
huggingface.co

Total runs: 11.6K
Run Growth: 11.0K
Growth Rate: 95.53%
Updated:August 07 2026
huggingface.co

Total runs: 10.8K
Run Growth: -659
Growth Rate: -6.11%
Updated:July 25 2026
huggingface.co

Total runs: 5.1K
Run Growth: 3.2K
Growth Rate: 62.77%
Updated:January 20 2026
huggingface.co

Total runs: 1.2K
Run Growth: -1.3K
Growth Rate: -113.47%
Updated:July 17 2025
huggingface.co

Total runs: 806
Run Growth: -1.5K
Growth Rate: -192.18%
Updated:August 29 2025
huggingface.co

Total runs: 642
Run Growth: -354
Growth Rate: -55.14%
Updated:June 12 2025
huggingface.co

Total runs: 484
Run Growth: 101
Growth Rate: 20.87%
Updated:July 24 2025
huggingface.co

Total runs: 366
Run Growth: 230
Growth Rate: 62.84%
Updated:July 17 2025
huggingface.co

Total runs: 226
Run Growth: -132
Growth Rate: -58.41%
Updated:August 11 2026
huggingface.co

Total runs: 129
Run Growth: -440
Growth Rate: -341.09%
Updated:August 19 2026
huggingface.co

Total runs: 52
Run Growth: -27
Growth Rate: -51.92%
Updated:July 29 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 30 2026