Efficient audio understanding with general audio captions
This repository contains the
fp8 quantized
weights of the original model, which provides substantial memory savings and faster inference throughput while retaining overall task performance close to the
bf16 release
. As quantization introduces numerical approximations, individual outputs may differ slightly from the full-precision model. If you need maximum numerical fidelity (e.g., strict reproduction), use the
fp32 model
.
🔥 Key Highlights
State-of-the-Art Performance
Outperforms Qwen2.5-Omni-7B, Kimi-Audio-Instruct-7B on
multiple key audio understanding tasks
.
High Efficiency
3.2×
throughput speedup at comparable batch sizes compared to Qwen2.5-Omni-7B.
20x
throughput speedup by increasing furhter batchsizes. We tested up to a
batch size=512
for 30s audio input on 80GB GPUs. Baselines only support batch size = 8.
Time-to-first-token (TTFT) speedup of up to
4x
compared to Qwen2.5-Omni-7B.
Caption-based Alignment
Trained with
general audio captions
(instead of ASR transcripts) to achieve holistic audio understanding.
Full Transparency
Public-source
training data and reproducible pipeline.
Apache License 2.0 for
both research and commercial use
.
Acknowledgment and Model Foundation
Although MiDashengLM demonstrates superior audio understanding performance and efficiency compared to Qwen2.5-Omni models,
we acknowledge
Qwen2.5-Omni as a remarkable and respected foundational work
in the field.
Our model specifically uses
Qwen2.5-Omni-7B Thinker
as the initialization for decoder training, building upon its robust architecture and weight initialization.
The audio encoder is built upon
Dasheng
, an open-source audio encoder for general audio understanding with state-of-the-art performance.
Dasheng serves as the core foundation enabling MiDashengLM's exceptional performance
.
Framework
MiDashengLM integrates the powerful Dasheng audio encoder with
the Qwen2.5-Omni-7B Thinker decoder through a unique caption-based alignment strategy.
Unlike conventional ASR-driven approaches,
our model leverages general audio captions to capture comprehensive audio representations encompassing speech, environmental sounds, and musical elements
in a unified textual format. This design enables holistic audio understanding while maintaining exceptional computational efficiency.
Why Captions Instead of ASR?
ASR Limitations:
Discards huge amount of non-speech audio (music/environmental sounds).
Misses paralinguistic info (speaker emotion, acoustic properties).
Non-monotonic alignment provides a hard learning signal.
Novel Open Source Dataset for Training: ACAVCaps
ACAVCaps is a meticulously curated 38,662-hour collection of general audio captions derived from the open-source
ACAV100M audio repository
.
While leveraging ACAV100M's extensive raw audio materials, we completely re-engineered the annotation process to create a dataset for holistic audio understanding.
We devide the dataset into six categories:
Category
Example Caption
Pure Speech
"A female voice narrates historical competition with synthetic modulation"
Pure Sound
"Outdoor scene with wind, birds, duck quacking and background noise"
Pure Music
"Crowd cheering with electronic synthesizer-driven soundscape"
Mixed Music
"The audio features a crowd cheering and clapping alongside electronic music with a synthesizer-driven, dark, and energetic soundscape."
Mixed Speech
"A Russian voice demonstrates a synthesizer’s capabilities over an experimental electronic backdrop, explaining its sound design and value in a gritty, vocal-fry tone."
Mixed Sound
"A man speaks in English about entering a city and village, accompanied by the sounds of a running vehicle."
The figure below illustrates our data curation pipeline for ACAVCaps:
Each caption is generated through a three-step process:
Please refer to the official repositories for evaluation on the
MECAT
and
MMAU
benchmarks.
Efficiency
MiDashengLM demonstrates superior inference efficiency compared to Qwen2.5-Omni-7B,
achieving 3.2× speedup at comparable batch sizes and an overall potential speedup of 20.2× with larger batches.
Batch Size
MiDashengLM (samples/s)
Qwen2.5-Omni-7B (samples/s)
Speedup
1
0.45
0.36
1.25x
4
1.40
0.91
1.53x
8
2.72
1.15
2.36x
16
5.18
OOM
-
32
9.78
OOM
-
64
17.07
OOM
-
128
22.73
OOM
-
200
25.15
OOM
-
Tested on 80GB GPU with 30s audio, 100-token output.
Training Data
MiDashengLM is trained exclusively on publicly available datasets across five categories: Speech, Sound and General Audio, Speech and Paralinguistic, Music, and Question Answering. All datasets are listed below with their respective tasks, lengths, and supervised fine-tuning (SFT) usage.
Speech Training Data
This table lists speech-related datasets used for tasks like Automatic Speech Recognition (ASR), keyword spotting (KWS), and speech-to-text translation (S2TT).
The column “SFT?” indicates whether the dataset is used for supervised fine-tuning.
Data
Task
Length(h)
SFT?
LibriSpeech
ASR
960
√
LibriHeavy
ASR
50,000
X
GigaSpeech
ASR
10,000
√
GigaSpeech2
ASR
30,000
√
WeNetSpeech
ASR
10,000
√
Yodas
ASR
320,000
X
CommonVoice-17.0
ASR
5,000
√
AISHELL-1
ASR
100
√
AISHELL-2
ASR
1,000
√
AISHELL-3
ASR
70
√
LJSpeech-1.1
ASR
37
X
LibriTTS
ASR
585
X
MultiLingualSpokenWords
KWS
5,000
X
Emilia
ASR
101,000
√
CovoST-v2
S2TT
2,880
√
Fleurs
S2TT
1,224
X
MSR-86K
ASR, LangID
86,000
√
ACAV100M-Speech
ASR
55,754
X
Must-C
ASR,S2TT
1,000
√
MLS
ASR
50,000
X
SpgiSpeech
ASR
5,000
X
PeoplesSpeech
ASR
30,000
X
KeSpeech
ASR
1,400
√
LAION-300M
Caption
230,000
X
Total
997,010
258.410
Sound and General Audio Datasets
Dataset
Task
Length(h)
SFT?
FSD50k
Sound Event
77
√
AudioSet
Sound Event
5,200
AudioSet-strong
Sound Event
220
X
VGGSound
Sound Event
540
√
FSDKaggle2018
Sound Event
20
√
FSDKaggle2019
Sound Event
100
ARCA23k
Sound Event
120
X
AutoACD
Audio(Sound) Caption
5,200
√
AudioSetCaps
Audio(Sound) Caption
6,000
√
SoundVECaps
Audio(Sound) Caption
5,000
√
WavCaps
Audio(Sound) Caption
7,567
√
Audiocaps
Audio(Sound) Caption
100
√
Clothov2
Audio(Sound) Caption
17
√
TACOS
Audio(Sound) Caption
98
√
CochlScene
SoundScape
500
√
BirdSet
SoundScape
7,000
X
ACAVCaps
General Caption
38,662
√
Total
76.421
69.081
Speech and Paralinguistic Datasets
Dataset
Task
Length(hours)
SFT?
IEMOCAP
Emotion
8
√
Meld
Emotion
12
√
SUBESCO
Emotion
9
X
RAVDESS-Speech
Emotion
2
X
RAVDESS-Song
Emotion
1
X
CREMA-D
Emotion
4
X
ESD
Emotion
29
X
VocalSound
Vocal sound classification
20
√
NonSpeech7k
Vocal sound classification
3
√
VoxLingua107
Language identification
7,200
√
CommonLanguage
Language identification
45
√
YLACombe
Language identification
5
X
VoxCeleb1
Speaker verification
76
√
CNCeleb
Speaker verification & age
2,100
√
VoxCeleb2
Speaker verification
1,000
√
VoxBlink1
Speaker verification
1,300
VoxBlink2
Speaker verification
2,600
√
VoxTube
Language identification
5,200
√
LibriCount
Speaker counting
8
√
FluentSpeechCommands
Intent classification & gender
17
X
SpeechOcean762
Speaker age
5
X
ASVSpoof5
Spoof detection
603
X
Total
20,247
19,572
Music-Related Datasets
Covers music captioning, genre recognition, instrument classification, and singing style identification.
Dataset
Task
Length(h)
SFT?
MusicCaps
Music Caption
15
√
Songdescriber
Music Caption
23
√
LPMusicCaps-MTT
Music Caption
18
√
LPMusicCaps-MSD
Music Caption
1,000
√
VocalSet
Singing style identification
10
X
FreeMusicArchive
Genre recognition
610
√
MTG-Jamendo
Instrument classification Genre recognition
3,768
√
NSynth
Instrument classification
360
√
GoodSounds
Instrument classification
28
√
chMusic
Instrument classification
1
√
CTIS
Instrument classification
1
√
Total
5,824
5,814
Question Answering Datasets
Used for training on audio-visual QA, environment QA, and music QA tasks. Most support SFT.
Dataset
Task
# QA
SFT?
AVQA
Environment QA
36,114
√
ClothoAQA
Environment QA
6,175
√
TACOS+
Environment QA
40,019
√
MusicQA
Music QA
112,878
√
SIFT-50M
Speech QA
21,430,000
√
ACAV-QA
General QA
24,371
√
Citation
MiDashengLM is under the Apache License 2.0, and we encourage its use in
both research and business applications
.
If you find MiDashengLM useful in your research, please consider citing our work:
@techreport{midashenglm7b,
title = {MiDashengLM: Efficient Audio Understanding with General Audio Captions},
author = {{Horizon Team, MiLM Plus}},
institution= {Xiaomi Inc.},
year = {2025},
note = {Contributors: Heinrich Dinkel et al. (listed alphabetically in Appendix B)},
url = {https://arxiv.org/abs/2508.03983},
eprint = {2508.03983},
}
Runs of mispeech midashenglm-7b-fp8 on huggingface.co
10
Total runs
0
24-hour runs
6
3-day runs
6
7-day runs
10
30-day runs
More Information About midashenglm-7b-fp8 huggingface.co Model
midashenglm-7b-fp8 huggingface.co is an AI model on huggingface.co that provides midashenglm-7b-fp8's model effect (), which can be used instantly with this mispeech midashenglm-7b-fp8 model. huggingface.co supports a free trial of the midashenglm-7b-fp8 model, and also provides paid use of the midashenglm-7b-fp8. Support call midashenglm-7b-fp8 model through api, including Node.js, Python, http.
midashenglm-7b-fp8 huggingface.co is an online trial and call api platform, which integrates midashenglm-7b-fp8's modeling effects, including api services, and provides a free online trial of midashenglm-7b-fp8, you can try midashenglm-7b-fp8 online for free by clicking the link below.
mispeech midashenglm-7b-fp8 online free url in huggingface.co:
midashenglm-7b-fp8 is an open source model from GitHub that offers a free installation service, and any user can find midashenglm-7b-fp8 on GitHub to install. At the same time, huggingface.co provides the effect of midashenglm-7b-fp8 install, users can directly use midashenglm-7b-fp8 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.