stepfun-ai / Step-Audio-AQAA

huggingface.co
Total runs: 218
24-hour runs: 7
7-day runs: 70
30-day runs: 153
Model's Last Updated: June 12 2025

Introduction of Step-Audio-AQAA

Model Details of Step-Audio-AQAA

Step-Audio-AQAA: A Fully End-to-End Expressive Large Audio Language Model

📚 Paper: Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

🚀 Live Demo: Try the Demo

Model Overview

Step-Audio-AQAA is a fully end-to-end Large Audio-Language Model (LALM) designed for Audio Query-Audio Answer (AQAA) tasks. It directly processes audio inputs and generates natural, accurate speech responses without relying on traditional ASR and TTS modules, eliminating cascading errors and simplifying the system architecture.

Key Capabilities
  • Fully End-to-End Audio Interaction : Generates speech outputs directly from raw audio inputs without ASR/TTS intermediates.
  • Fine-Grained Voice Control : Supports sentence-level adjustments of emotional tone, speech rate, and other vocal features.
  • Multilingual & Dialect Support : Covers Chinese (including Sichuanese, Cantonese), English, Japanese, etc.
  • Complex Task Handling : Excels in speech emotion control, role-playing, logical reasoning, and other complex audio interactions.
Model Architecture

Step-Audio-AQAA consists of three core modules:

Dual-Codebook Audio Tokenizer
  • Linguistic Tokenizer : Based on Paraformer encoder, extracts phonemic and linguistic attributes with a 1,024-codebook at 16.7Hz.
  • Semantic Tokenizer : References CosyVoice 1.0, captures acoustic features with a 4,096-codebook at 25Hz.
  • Temporal Alignment : Uses a 2:3 interleaving ratio to ensure temporal consistency between token types.
Backbone LLM
  • Parameter Scale : 130-billion-parameter multi-modal LLM (Step-Omni).
  • Architecture : Decoder-only with Transformer blocks, RMSNorm layers, and grouped query attention.
  • Vocabulary Expansion : Incorporates 5,120 audio tokens into the text vocabulary for text-audio interleaved output.
Neural Vocoder
  • Architecture : Flow-matching model based on CosyVoice, using U-Net and ResNet-1D layers.
  • Conditional Generation : Generates high-fidelity speech waveforms conditioned solely on audio tokens.
Training Approach
Multi-Stage Training Pipeline
  1. Pretraining : Multi-modal pretraining on text, audio, and image data.
  2. Supervised Fine-Tuning (SFT) :
    • Stage 1: Full-parameter update on AQTA and AQTAA datasets.
    • Stage 2: Optimizes specific capabilities with high-quality AQTAA data.
  3. Direct Preference Optimization (DPO) : Uses audio token masking to avoid degradation of speech generation.
  4. Model Merging : Weighted combination of SFT and DPO models to enhance overall performance.
Training Data
  • Multi-Modal Pretraining Data : 800 billion text tokens and audio-text interleaved data.
  • AQTA Dataset : Audio query-text answer pairs.
  • AQTAA Dataset : Audio query-text answer-audio answer triplets generated from AQTA.
Citation
@misc{huang2025stepaudioaqaa,
      title={Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model}, 
      author={Ailin Huang and Boyong Wu and Bruce Wang and Chao Yan and Chen Hu and Chengli Feng and Fei Tian and Feiyu Shen and Jingbei Li and Mingrui Chen and et al.},
      year={2025},
      eprint={2506.08967},
      archivePrefix={arXiv},
      primaryClass={cs.SD}
}
Team & Contributions

Step-Audio-AQAA is developed by the StepFun team, with contributions from multiple researchers and engineers. For technical support or collaboration, contact the corresponding authors: Daxin Jiang ( [email protected] ), Shuchang Zhou ( [email protected] ), Chen Hu ( [email protected] ).

License

This model is released under the Apache 2.0 license. For more details, please refer to the license file.

Runs of stepfun-ai Step-Audio-AQAA on huggingface.co

218
Total runs
7
24-hour runs
25
3-day runs
70
7-day runs
153
30-day runs

More Information About Step-Audio-AQAA huggingface.co Model

More Step-Audio-AQAA license Visit here:

https://choosealicense.com/licenses/apache-2.0

Step-Audio-AQAA huggingface.co

Step-Audio-AQAA huggingface.co is an AI model on huggingface.co that provides Step-Audio-AQAA's model effect (), which can be used instantly with this stepfun-ai Step-Audio-AQAA model. huggingface.co supports a free trial of the Step-Audio-AQAA model, and also provides paid use of the Step-Audio-AQAA. Support call Step-Audio-AQAA model through api, including Node.js, Python, http.

Step-Audio-AQAA huggingface.co Url

https://huggingface.co/stepfun-ai/Step-Audio-AQAA

stepfun-ai Step-Audio-AQAA online free

Step-Audio-AQAA huggingface.co is an online trial and call api platform, which integrates Step-Audio-AQAA's modeling effects, including api services, and provides a free online trial of Step-Audio-AQAA, you can try Step-Audio-AQAA online for free by clicking the link below.

stepfun-ai Step-Audio-AQAA online free url in huggingface.co:

https://huggingface.co/stepfun-ai/Step-Audio-AQAA

Step-Audio-AQAA install

Step-Audio-AQAA is an open source model from GitHub that offers a free installation service, and any user can find Step-Audio-AQAA on GitHub to install. At the same time, huggingface.co provides the effect of Step-Audio-AQAA install, users can directly use Step-Audio-AQAA installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Step-Audio-AQAA install url in huggingface.co:

https://huggingface.co/stepfun-ai/Step-Audio-AQAA

Url of Step-Audio-AQAA

Step-Audio-AQAA huggingface.co Url

Provider of Step-Audio-AQAA huggingface.co

stepfun-ai
ORGANIZATIONS

Other API from stepfun-ai

huggingface.co

Total runs: 605.4K
Run Growth: -87.7K
Growth Rate: -14.51%
Updated:February 04 2025
huggingface.co

Total runs: 41.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 29 2026
huggingface.co

Total runs: 4.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 02 2025
huggingface.co

Total runs: 83
Run Growth: -209
Growth Rate: -251.81%
Updated:January 14 2026