skt / A.X-K2-ALM

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: July 30 2026
audio-to-audio

Introduction of A.X-K2-ALM

Model Details of A.X-K2-ALM

A.X K2 ALM โ€” Audio Language Model

A.X K2 ALM

๐Ÿค— Models | ๐Ÿ“„ Technical Report

Model Summary

A.X K2 ALM (Audio Language Model) is a member of the A.X K2 family that adds a speech modality to a frozen in-house LLM backbone โ€” A.X K2 Light (20B-A2.7B MoE) โ€” to enable natural voice-based conversation. It integrates speech understanding, voice activity detection (VAD), and speech generation within a single unified, streaming architecture built from four components: a chunk-wise Conformer speech encoder, an adapter (MLP projector + Q-Former), a lightweight VAD module, and a multi-token-prediction (MTP) speech decoder.

Motivated by production settings where the deployed LLM may be audio-unaware and its existing text-response quality must be preserved, we keep the LLM entirely frozen and train only the speech-side components โ€” avoiding the catastrophic forgetting and text-quality degradation commonly induced by end-to-end multimodal training, and letting the same recipe transfer readily to other LLMs. On our Korean KVoiceBench , A.X K2 ALM reaches 95% of Qwen3-Omni-30B-A3B's response quality at roughly two-thirds the parameters .

A.X K2 ALM overall architecture
A.X K2 ALM architecture. Speech encoder โ†’ adapter โ†’ frozen LLM โ†’ speech decoder; the VAD reads the adapter and LLM representations.
Key Features
  • A Korean ALM Built Entirely from Scratch Every speech-side component โ€” the streaming speech encoder, the adapter, the VAD heads, and the speech decoder with its neural audio codec โ€” is built and trained in-house from scratch rather than adapted from open-source models, on top of the frozen A.X K2 Light backbone. This lets the frame rate, streaming behavior, and language coverage be tuned end-to-end for Korean voice services.

  • Self-Supervised (SSL) Streaming Speech Encoder A 300M chunk-wise Conformer encoder, self-supervised pre-trained from scratch on 400K hours of unlabeled speech with multi-codebook Random Quantization and fine-tuned as a 341M streaming transducer, outperforms Qwen3-Omni on Korean ASR benchmarks at roughly half its audio-encoder size โ€” even under a low-latency streaming constraint.

  • Efficient Streaming Speech Decoder We introduce a 0.88B multi-token speech decoder that integrates a textโ€“speech language model with a depth-transformer to hierarchically generate semantic tokens and residual acoustic codes in real time. Operating at roughly half the parameter footprint of current SOTA baselines, our model achieves superior pronunciation accuracy with the lowest CER while maintaining high speech intelligibility comparable to leading systems, ensuring precise delivery of generated text responses in real-time ALMs.

  • Forgetting-Free Modality Extension The LLM stays fully frozen; a two-stage adapter โ€” speech alignment (syllable-CTC / char2token bridge) then self-distillation dialogue training โ€” transfers the backbone's response behavior to the audio path without changing its text-input behavior.

  • Integrated Context-Aware VAD The external VAD is replaced by lightweight six-class heads on the model's own representations โ€” an adapter-based head detects onset, an LLM-based head verifies with context โ€” enabling barge-in and far fewer false triggers on non-speech and mid-utterance pauses.

Model Details
Function Component Configuration Params
Speech Understanding Speech Encoder Chunk-wise Conformer 300M
Adapter โ€” MLP Projector 2-layer, GELU 5.77M
Adapter โ€” char-CTC head linear, 2048 โ†’ 2667 5.46M
Adapter โ€” Q-Former 4L / 16H / d=2048 / FFN 8192 201.4M
LLM (frozen) A.X K2 Light (in-house MoE) 20B-A2.7B
Speech Generation Conditioning Projector FFN stack (2048 โ†’ 1024, 2 paths) 71.3M
Multi-Token Predictor Textโ€“Speech Language Model 630M
Depth Transformer 110M
Speaker Encoder Transformer 100M
Audio Codec Decoder Transformer + ConvNet 40M
Total A.X K2 ALM โ‰ˆ22B-A4B
  • VAD heads: six-class classifiers over the adapter output and LLM hidden states; parameter counts are negligible.
  • Languages: Korean
  • Streaming: all input/output paths are designed for streaming operation for low-latency live voice interaction.
  • Full details: see the technical report for the architecture and training recipe.
Evaluation Results

Speech understanding is evaluated on KVoiceBench , our own Korean spoken-query benchmark suite built in-house by localizing VoiceBench with human review and adding Korean-native professional-knowledge items, covering speech recognition (STT), spoken QA (Speech-QA), and instruction following (Speech-IF). Open-ended (OPEN) categories are scored by an LLM judge and knowledge (QA) categories by correctness. See the technical report for KVoiceBench construction details. Bold = best among compared models.

Task Benchmark Metric A.X K2 ALM (22B-A4B) Qwen2.5-Omni (7B) Qwen3-Omni (30B-A3B) HyperCLOVA X 8B Omni
STT KsponSpeech (eval-clean) CER โ†“ 9.00 18.96 8.46 10.22
KsponSpeech (eval-other) CER โ†“ 9.12 22.72 7.91 10.15
Speech-QA AdvBench OPEN โ†‘ 92.67 92.61 89.55 81.87
AlpacaEval OPEN โ†‘ 69.70 52.87 71.14 51.81
AlpacaEval-Full OPEN โ†‘ 70.75 52.45 72.21 52.43
BBH QA โ†‘ 44.18 52.66 57.22 43.80
CommonEval OPEN โ†‘ 70.20 55.41 68.25 55.49
MMSU QA โ†‘ 43.67 32.08 43.48 32.08
OpenBookQA QA โ†‘ 76.62 76.12 91.54 70.90
SD-QA QA โ†‘ 46.87 25.89 48.23 27.52
WildVoice OPEN โ†‘ 66.88 46.33 66.25 45.77
Speech-IF IFEval OPEN โ†‘ 61.20 44.88 67.61 35.21
Avg. Speech-QA / IF โ†‘ 64.27 53.13 67.55 49.69
Speech Generation KVoiceBench (Subset-200) CER โ†“ 1.86 โ€“ 26.37 4.16
UTMOS โ†‘ 2.871 โ€“ 3.178 3.171
STOI โ†‘ 0.990 โ€“ 0.999 0.996

Averaged over the ten KVoiceBench subsets, A.X K2 ALM scores 64.27 โ€” 95% of Qwen3-Omni-30B-A3B (67.55) at roughly two-thirds of its parameters โ€” and outperforms Qwen2.5-Omni-7B and HyperCLOVA X 8B Omni by 21% and 29% (relative).

Demos

1. Natural Voice Conversation A fluid, natural spoken dialogue with the model.

2. Context-Aware Turn-Taking The context-aware VAD ignores listener back-channels (e.g., "์Œ", "์•„") and waits through mid-sentence hesitations โ€” so it neither falsely stops nor cuts in, responding only once the user has actually finished.

Availability

Model weights are planned for public release (coming soon) . For now, this repository serves as the model card and hosts the demo videos above and the technical report for reference.

License

Released under the h-research license.

Citation

If you use A.X K2 ALM in your research, please cite the technical report:

@techreport{axk2alm2026,
  title       = {A.X K2 ALM (Audio Language Model)},
  author      = {SKT A.X Team},
  year        = {2026},
  institution = {SK Telecom}
}
Contact

For A.X models for business applications or scalable deployment, please contact [email protected] .

Runs of skt A.X-K2-ALM on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About A.X-K2-ALM huggingface.co Model

More A.X-K2-ALM license Visit here:

https://choosealicense.com/licenses/h-research

A.X-K2-ALM huggingface.co

A.X-K2-ALM huggingface.co is an AI model on huggingface.co that provides A.X-K2-ALM's model effect (), which can be used instantly with this skt A.X-K2-ALM model. huggingface.co supports a free trial of the A.X-K2-ALM model, and also provides paid use of the A.X-K2-ALM. Support call A.X-K2-ALM model through api, including Node.js, Python, http.

A.X-K2-ALM huggingface.co Url

https://huggingface.co/skt/A.X-K2-ALM

skt A.X-K2-ALM online free

A.X-K2-ALM huggingface.co is an online trial and call api platform, which integrates A.X-K2-ALM's modeling effects, including api services, and provides a free online trial of A.X-K2-ALM, you can try A.X-K2-ALM online for free by clicking the link below.

skt A.X-K2-ALM online free url in huggingface.co:

https://huggingface.co/skt/A.X-K2-ALM

A.X-K2-ALM install

A.X-K2-ALM is an open source model from GitHub that offers a free installation service, and any user can find A.X-K2-ALM on GitHub to install. At the same time, huggingface.co provides the effect of A.X-K2-ALM install, users can directly use A.X-K2-ALM installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

A.X-K2-ALM install url in huggingface.co:

https://huggingface.co/skt/A.X-K2-ALM

Url of A.X-K2-ALM

A.X-K2-ALM huggingface.co Url

Provider of A.X-K2-ALM huggingface.co

skt
ORGANIZATIONS

Other API from skt

huggingface.co

Total runs: 312.0K
Run Growth: -117.1K
Growth Rate: -37.52%
Updated:September 24 2021
huggingface.co

Total runs: 207.0K
Run Growth: 202.4K
Growth Rate: 99.19%
Updated:August 07 2026
huggingface.co

Total runs: 41.8K
Run Growth: -55.5K
Growth Rate: -127.35%
Updated:September 01 2026
huggingface.co

Total runs: 13.9K
Run Growth: 62
Growth Rate: 0.45%
Updated:July 23 2025
huggingface.co

Total runs: 12.6K
Run Growth: 11.9K
Growth Rate: 95.84%
Updated:August 07 2026
huggingface.co

Total runs: 11.9K
Run Growth: -350
Growth Rate: -2.95%
Updated:June 12 2025
huggingface.co

Total runs: 10.9K
Run Growth: -588
Growth Rate: -5.42%
Updated:July 25 2026
huggingface.co

Total runs: 5.6K
Run Growth: 3.4K
Growth Rate: 63.82%
Updated:January 20 2026
huggingface.co

Total runs: 1.2K
Run Growth: -1.3K
Growth Rate: -113.47%
Updated:July 17 2025
huggingface.co

Total runs: 806
Run Growth: -1.5K
Growth Rate: -192.18%
Updated:August 29 2025
huggingface.co

Total runs: 642
Run Growth: -354
Growth Rate: -55.14%
Updated:June 12 2025
huggingface.co

Total runs: 484
Run Growth: 101
Growth Rate: 20.87%
Updated:July 24 2025
huggingface.co

Total runs: 366
Run Growth: 230
Growth Rate: 62.84%
Updated:July 17 2025
huggingface.co

Total runs: 231
Run Growth: -129
Growth Rate: -56.33%
Updated:August 11 2026
huggingface.co

Total runs: 126
Run Growth: -442
Growth Rate: -348.03%
Updated:August 19 2026
huggingface.co

Total runs: 54
Run Growth: -26
Growth Rate: -49.06%
Updated:July 29 2026