A.X K2
is a large-scale Mixture-of-Experts (MoE) language model trained from scratch as a high-performance,
agentic
foundation model, and the successor to A.X K1.
The model contains
688 billion total parameters
, with
33 billion active parameters
, delivering strong reasoning and instruction-following performance while maintaining practical inference efficiency.
Through a
Think-Fusion
training recipe, a single unified model supports both a
thinking
mode for complex problem solving and a
non-thinking
mode for concise, low-latency responses, allowing the user to trade quality for cost on a per-request basis.
A.X K2 is developed as part of the Korean government's
Sovereign AI foundation model
project, aiming to build a frontier-scale model with deep understanding of the Korean language and culture.
A.X K2 NVFP4
A.X K2 NVFP4 is an NVFP4-quantized version of
A.X K2
, a 688B-parameter Mixture-of-Experts language model developed by SK Telecom.
This checkpoint applies
NVFP4 W4A4 quantization to the routed experts
, while keeping the remaining modules in FP8 or BF16. It retains performance comparable to the FP8 checkpoint while reducing the model memory footprint by approximately half.
The model can be served on a single node with
4 NVIDIA B200 GPUs
using the A.X K2–enabled SKT-AI vLLM fork.
Set
enable_thinking
to
false
for a direct non-thinking response.
Long Context
The checkpoint includes the YaRN configuration required for context lengths of up to
256K tokens
. No separate RoPE override is required when using the provided
config.json
.
The serving context can be reduced with
--max-model-len
when longer inputs are not required.
Needle-in-a-Haystack (NIAH)
NIAH probes exact fact retrieval by inserting a target fact (the "needle") at varying depths within a long context and asking the model to recover it. Under zero-shot YaRN scaling (scaling factors of 1, 2, and 4 for 128K, 256K, and 512K, respectively), A.X K2 attains a perfect retrieval score at every context length and needle depth—and does so even after NVFP4 (experts-only W4A4) quantization. The heatmaps below show the 256K and 512K results.
NIAH fact retrieval across context length (x-axis) and needle depth (y-axis): 256K (YaRN factor 2, top) and 512K (YaRN factor 4, bottom), under NVFP4 (experts-only W4A4) quantization. A.X K2 scores a perfect 100 at every position.
A.X-K2-NVFP4 huggingface.co is an AI model on huggingface.co that provides A.X-K2-NVFP4's model effect (), which can be used instantly with this skt A.X-K2-NVFP4 model. huggingface.co supports a free trial of the A.X-K2-NVFP4 model, and also provides paid use of the A.X-K2-NVFP4. Support call A.X-K2-NVFP4 model through api, including Node.js, Python, http.
A.X-K2-NVFP4 huggingface.co is an online trial and call api platform, which integrates A.X-K2-NVFP4's modeling effects, including api services, and provides a free online trial of A.X-K2-NVFP4, you can try A.X-K2-NVFP4 online for free by clicking the link below.
skt A.X-K2-NVFP4 online free url in huggingface.co:
A.X-K2-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find A.X-K2-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of A.X-K2-NVFP4 install, users can directly use A.X-K2-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.