mispeech / r1-aqa

huggingface.co
Total runs: 70
24-hour runs: 0
7-day runs: 21
30-day runs: -17
Model's Last Updated: March 29 2025
audio-text-to-text

Introduction of r1-aqa

Model Details of r1-aqa

R1-AQA --- Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Introduction

R1-AQA is a audio question answering (AQA) model based on Qwen2-Audio-7B-Instruct , optimized through reinforcement learning using the group relative policy optimization (GRPO) algorithm. This implementation has achieved state-of-the-art performance on MMAU Test-mini benchmark with only 38k post-training samples. For more details, please refer to our Github and Technical Report .

Table: Accuracies (%) on MMAU Test-mini benchmark
Model Method Sound Music Speech Average
\ Human* 86.31 78.22 82.17 82.23
Gemini Pro 2.0 Flash Direct Inference* 56.46 58.68 51.65 55.60
Audio Flamingo 2 Direct Inference* 61.56 73.95 30.93 55.48
GPT4o + Strong Cap. Direct Inference* 57.35 49.70 64.86 57.30
Llama-3-8B-Instruct + Strong Cap. Direct Inference* 50.75 48.93 55.25 52.10
Gemini Pro v1.5 Direct Inference* 56.75 49.40 58.55 54.90
Qwen2-Audio-7B-Instruct Direct Inference* 54.95 50.98 42.04 49.20
GPT4o + Weak Cap. Direct Inference* 39.33 41.90 58.25 45.70
Llama-3-8B-Instruct + Weak Cap. Direct Inference* 34.23 38.02 54.05 42.10
SALMONN Direct Inference* 41.00 34.80 25.50 33.70
Qwen2-Audio-7B-Instruct CoTA [1] 60.06 64.30 60.70 61.71
Qwen2-Audio-7B-Instruct Zero-Shot-CoT [2] 61.86 56.29 55.26 57.80
Qwen2-Audio-7B-Instruct GRPO (Ours) 69.37 66.77 57.36 64.50
Notes:

* The data are sourced from the MMAU official website: https://sakshi113.github.io/mmau_homepage/
[1] Xie, Zhifei, et al. "Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models." arXiv preprint arXiv:2503.02318 (2025).
[2] Ma, Ziyang, et al. "Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model." arXiv preprint arXiv:2501.07246 (2025).

Inference
import torch
import torchaudio
from transformers import Qwen2AudioForConditionalGeneration, AutoProcessor

# Load model
model_name = "mispeech/r1-aqa"
processor = AutoProcessor.from_pretrained(model_name)
model = Qwen2AudioForConditionalGeneration.from_pretrained(model_name, torch_dtype=torch.bfloat16, device_map="auto")

# Load example audio
wav_path = "test-mini-audios/3fe64f3d-282c-4bc8-a753-68f8f6c35652.wav"  # from MMAU dataset
waveform, sampling_rate = torchaudio.load(wav_path)
assert sampling_rate == 16000
audios = [waveform.numpy()]

# Make prompt text
question = "Based on the given audio, identify the source of the speaking voice."
options = ["Man", "Woman", "Child", "Robot"]
prompt = f"{question} Please choose the answer from the following options: {str(options)}. Output the final answer in <answer> </answer>."
message = [
    {"role": "user", "content": [
        {"type": "audio", "audio_url": wav_path},
        {"type": "text", "text": prompt}
    ]}
]
texts = processor.apply_chat_template(message, add_generation_prompt=True, tokenize=False)

# Process
inputs = processor(text=texts, audios=audios, sampling_rate=16000, return_tensors="pt", padding=True).to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=256)
generated_ids = generated_ids[:, inputs.input_ids.size(1):]
response = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)

print(response)

Runs of mispeech r1-aqa on huggingface.co

70
Total runs
0
24-hour runs
1
3-day runs
21
7-day runs
-17
30-day runs

More Information About r1-aqa huggingface.co Model

r1-aqa huggingface.co

r1-aqa huggingface.co is an AI model on huggingface.co that provides r1-aqa's model effect (), which can be used instantly with this mispeech r1-aqa model. huggingface.co supports a free trial of the r1-aqa model, and also provides paid use of the r1-aqa. Support call r1-aqa model through api, including Node.js, Python, http.

mispeech r1-aqa online free

r1-aqa huggingface.co is an online trial and call api platform, which integrates r1-aqa's modeling effects, including api services, and provides a free online trial of r1-aqa, you can try r1-aqa online for free by clicking the link below.

mispeech r1-aqa online free url in huggingface.co:

https://huggingface.co/mispeech/r1-aqa

r1-aqa install

r1-aqa is an open source model from GitHub that offers a free installation service, and any user can find r1-aqa on GitHub to install. At the same time, huggingface.co provides the effect of r1-aqa install, users can directly use r1-aqa installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

r1-aqa install url in huggingface.co:

https://huggingface.co/mispeech/r1-aqa

Url of r1-aqa

Provider of r1-aqa huggingface.co

mispeech
ORGANIZATIONS

Other API from mispeech

huggingface.co

Total runs: 21.0K
Run Growth: -4.3K
Growth Rate: -20.43%
Updated:March 30 2026
huggingface.co

Total runs: 4.6K
Run Growth: -2.6K
Growth Rate: -56.88%
Updated:March 19 2026
huggingface.co

Total runs: 4.5K
Run Growth: -2.3K
Growth Rate: -50.10%
Updated:March 30 2026
huggingface.co

Total runs: 4.1K
Run Growth: -33
Growth Rate: -0.78%
Updated:March 30 2026
huggingface.co

Total runs: 1.5K
Run Growth: -1.8K
Growth Rate: -104.32%
Updated:March 30 2026
huggingface.co

Total runs: 332
Run Growth: -490
Growth Rate: -147.15%
Updated:March 26 2026