We present Kimi-Audio, an open-source audio foundation model excelling in
audio understanding, generation, and conversation
. This repository hosts the model checkpoints for Kimi-Audio-7B.
Kimi-Audio is designed as a universal audio foundation model capable of handling a wide variety of audio processing tasks within a single unified framework. Key features include:
Universal Capabilities:
Handles diverse tasks like speech recognition (ASR), audio question answering (AQA), audio captioning (AAC), speech emotion recognition (SER), sound event/scene classification (SEC/ASC) and end-to-end speech conversation.
State-of-the-Art Performance:
Achieves SOTA results on numerous audio benchmarks (see our
Technical Report
).
Large-Scale Pre-training:
Pre-trained on over 13 million hours of diverse audio data (speech, music, sounds) and text data.
Novel Architecture:
Employs a hybrid audio input (continuous acoustic + discrete semantic tokens) and an LLM core with parallel heads for text and audio token generation.
Efficient Inference:
Features a chunk-wise streaming detokenizer based on flow matching for low-latency audio generation.
Kimi-Audio-7B is a base model without fine-tuning. So it cannot be used directly.
The base model is quite flexible, you can fine-tune it on any possible downstream tasks.
Kimi-Audio-7B huggingface.co is an AI model on huggingface.co that provides Kimi-Audio-7B's model effect (), which can be used instantly with this moonshotai Kimi-Audio-7B model. huggingface.co supports a free trial of the Kimi-Audio-7B model, and also provides paid use of the Kimi-Audio-7B. Support call Kimi-Audio-7B model through api, including Node.js, Python, http.
Kimi-Audio-7B huggingface.co is an online trial and call api platform, which integrates Kimi-Audio-7B's modeling effects, including api services, and provides a free online trial of Kimi-Audio-7B, you can try Kimi-Audio-7B online for free by clicking the link below.
moonshotai Kimi-Audio-7B online free url in huggingface.co:
Kimi-Audio-7B is an open source model from GitHub that offers a free installation service, and any user can find Kimi-Audio-7B on GitHub to install. At the same time, huggingface.co provides the effect of Kimi-Audio-7B install, users can directly use Kimi-Audio-7B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.