A.X K2 ALM
(Audio Language Model) is a member of the
A.X K2 family
that adds a speech modality to a frozen in-house LLM backbone โ
A.X K2 Light
(20B-A2.7B MoE) โ to enable natural voice-based conversation. It integrates
speech understanding, voice activity detection (VAD), and speech generation within a single unified, streaming architecture
built from four components: a chunk-wise Conformer speech encoder, an adapter (MLP projector + Q-Former), a lightweight VAD module, and a multi-token-prediction (MTP) speech decoder.
Motivated by production settings where the deployed LLM may be audio-unaware and its existing text-response quality must be preserved, we keep the LLM
entirely frozen
and train only the speech-side components โ avoiding the catastrophic forgetting and text-quality degradation commonly induced by end-to-end multimodal training, and letting the same recipe transfer readily to other LLMs. On our Korean
KVoiceBench
, A.X K2 ALM reaches
95% of Qwen3-Omni-30B-A3B's response quality at roughly two-thirds the parameters
.
A.X K2 ALM architecture.
Speech encoder โ adapter โ frozen LLM โ speech decoder; the VAD reads the adapter and LLM representations.
Key Features
A Korean ALM Built Entirely from Scratch
Every speech-side component โ the streaming speech encoder, the adapter, the VAD heads, and the speech decoder with its neural audio codec โ is built and trained
in-house from scratch
rather than adapted from open-source models, on top of the frozen A.X K2 Light backbone. This lets the frame rate, streaming behavior, and language coverage be tuned end-to-end for Korean voice services.
Self-Supervised (SSL) Streaming Speech Encoder
A 300M chunk-wise Conformer encoder, self-supervised pre-trained from scratch on 400K hours of unlabeled speech with multi-codebook Random Quantization and fine-tuned as a 341M streaming transducer, outperforms Qwen3-Omni on Korean ASR benchmarks at roughly half its audio-encoder size โ even under a low-latency streaming constraint.
Efficient Streaming Speech Decoder
We introduce a 0.88B multi-token speech decoder that integrates a textโspeech language model with a depth-transformer to hierarchically generate semantic tokens and residual acoustic codes in real time. Operating at roughly half the parameter footprint of current SOTA baselines, our model achieves superior pronunciation accuracy with the lowest CER while maintaining high speech intelligibility comparable to leading systems, ensuring precise delivery of generated text responses in real-time ALMs.
Forgetting-Free Modality Extension
The LLM stays fully frozen; a two-stage adapter โ speech alignment (syllable-CTC / char2token bridge) then self-distillation dialogue training โ transfers the backbone's response behavior to the audio path without changing its text-input behavior.
Integrated Context-Aware VAD
The external VAD is replaced by lightweight six-class heads on the model's own representations โ an adapter-based head detects onset, an LLM-based head verifies with context โ enabling
barge-in
and far fewer false triggers on non-speech and mid-utterance pauses.
Model Details
Function
Component
Configuration
Params
Speech Understanding
Speech Encoder
Chunk-wise Conformer
300M
Adapter โ MLP Projector
2-layer, GELU
5.77M
Adapter โ char-CTC head
linear, 2048 โ 2667
5.46M
Adapter โ Q-Former
4L / 16H / d=2048 / FFN 8192
201.4M
LLM (frozen)
A.X K2 Light (in-house MoE)
20B-A2.7B
Speech Generation
Conditioning Projector
FFN stack (2048 โ 1024, 2 paths)
71.3M
Multi-Token Predictor
TextโSpeech Language Model
630M
Depth Transformer
110M
Speaker Encoder
Transformer
100M
Audio Codec Decoder
Transformer + ConvNet
40M
Total
A.X K2 ALM
โ22B-A4B
VAD heads:
six-class classifiers over the adapter output and LLM hidden states; parameter counts are negligible.
Languages:
Korean
Streaming:
all input/output paths are designed for streaming operation for low-latency live voice interaction.
Full details:
see the
technical report
for the architecture and training recipe.
Evaluation Results
Speech understanding is evaluated on
KVoiceBench
, our own Korean spoken-query benchmark suite built in-house by localizing VoiceBench with human review and adding Korean-native professional-knowledge items, covering speech recognition (STT), spoken QA (Speech-QA), and instruction following (Speech-IF). Open-ended (OPEN) categories are scored by an LLM judge and knowledge (QA) categories by correctness. See the
technical report
for KVoiceBench construction details.
Bold
= best among compared models.
Task
Benchmark
Metric
A.X K2 ALM (22B-A4B)
Qwen2.5-Omni (7B)
Qwen3-Omni (30B-A3B)
HyperCLOVA X 8B Omni
STT
KsponSpeech (eval-clean)
CER โ
9.00
18.96
8.46
10.22
KsponSpeech (eval-other)
CER โ
9.12
22.72
7.91
10.15
Speech-QA
AdvBench
OPEN โ
92.67
92.61
89.55
81.87
AlpacaEval
OPEN โ
69.70
52.87
71.14
51.81
AlpacaEval-Full
OPEN โ
70.75
52.45
72.21
52.43
BBH
QA โ
44.18
52.66
57.22
43.80
CommonEval
OPEN โ
70.20
55.41
68.25
55.49
MMSU
QA โ
43.67
32.08
43.48
32.08
OpenBookQA
QA โ
76.62
76.12
91.54
70.90
SD-QA
QA โ
46.87
25.89
48.23
27.52
WildVoice
OPEN โ
66.88
46.33
66.25
45.77
Speech-IF
IFEval
OPEN โ
61.20
44.88
67.61
35.21
Avg.
Speech-QA / IF
โ
64.27
53.13
67.55
49.69
Speech Generation
KVoiceBench (Subset-200)
CER โ
1.86
โ
26.37
4.16
UTMOS โ
2.871
โ
3.178
3.171
STOI โ
0.990
โ
0.999
0.996
Averaged over the ten KVoiceBench subsets, A.X K2 ALM scores
64.27 โ 95% of Qwen3-Omni-30B-A3B (67.55) at roughly two-thirds of its parameters
โ and outperforms Qwen2.5-Omni-7B and HyperCLOVA X 8B Omni by 21% and 29% (relative).
Demos
1. Natural Voice Conversation
A fluid, natural spoken dialogue with the model.
2. Context-Aware Turn-Taking
The context-aware VAD ignores listener back-channels (e.g., "์", "์") and waits through mid-sentence hesitations โ so it neither falsely stops nor cuts in, responding only once the user has actually finished.
Availability
Model weights are
planned for public release (coming soon)
. For now, this repository serves as the model card and hosts the demo videos above and the
technical report
for reference.
License
Released under the
h-research
license.
Citation
If you use A.X K2 ALM in your research, please cite the technical report:
@techreport{axk2alm2026,
title = {A.X K2 ALM (Audio Language Model)},
author = {SKT A.X Team},
year = {2026},
institution = {SK Telecom}
}
Contact
For A.X models for business applications or scalable deployment, please contact
[email protected]
.
Runs of skt A.X-K2-ALM on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About A.X-K2-ALM huggingface.co Model
A.X-K2-ALM huggingface.co is an AI model on huggingface.co that provides A.X-K2-ALM's model effect (), which can be used instantly with this skt A.X-K2-ALM model. huggingface.co supports a free trial of the A.X-K2-ALM model, and also provides paid use of the A.X-K2-ALM. Support call A.X-K2-ALM model through api, including Node.js, Python, http.
A.X-K2-ALM huggingface.co is an online trial and call api platform, which integrates A.X-K2-ALM's modeling effects, including api services, and provides a free online trial of A.X-K2-ALM, you can try A.X-K2-ALM online for free by clicking the link below.
A.X-K2-ALM is an open source model from GitHub that offers a free installation service, and any user can find A.X-K2-ALM on GitHub to install. At the same time, huggingface.co provides the effect of A.X-K2-ALM install, users can directly use A.X-K2-ALM installed effect in huggingface.co for debugging and trial. It also supports api for free installation.