Qwen2-Audio is the new series of Qwen large audio-language models. Qwen2-Audio is capable of accepting various audio signal inputs and performing audio analysis or direct textual responses with regard to speech instructions. We introduce two distinct audio interaction modes:
voice chat: users can freely engage in voice interactions with Qwen2-Audio without text input;
audio analysis: users could provide audio and text instructions for analysis during the interaction;
We release Qwen2-Audio-7B and Qwen2-Audio-7B-Instruct, which are pretrained model and chat model respectively.
The code of Qwen2-Audio has been in the latest Hugging face transformers and we advise you to build from source with command
pip install git+https://github.com/huggingface/transformers
, or you might encounter the following error:
KeyError: 'qwen2-audio'
Quickstart
Here provides offers a code snippet illustrating the process of loading both the processor and model, alongside detailed instructions on executing the pretrained Qwen2-Audio base model for content generation.
from io import BytesIO
from urllib.request import urlopen
import librosa
from transformers import AutoProcessor, Qwen2AudioForConditionalGeneration
model = Qwen2AudioForConditionalGeneration.from_pretrained("Qwen/Qwen2-Audio-7B" ,trust_remote_code=True)
processor = AutoProcessor.from_pretrained("Qwen/Qwen2-Audio-7B" ,trust_remote_code=True)
prompt = "<|audio_bos|><|AUDIO|><|audio_eos|>Generate the caption in English:"
url = "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-Audio/glass-breaking-151256.mp3"
audio, sr = librosa.load(BytesIO(urlopen(url).read()), sr=processor.feature_extractor.sampling_rate)
inputs = processor(text=prompt, audios=audio, return_tensors="pt")
generated_ids = model.generate(**inputs, max_length=256)
generated_ids = generated_ids[:, inputs.input_ids.size(1):]
response = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]
Citation
If you find our work helpful, feel free to give us a cite.
@article{Qwen2-Audio,
title={Qwen2-Audio Technical Report},
author={Chu, Yunfei and Xu, Jin and Yang, Qian and Wei, Haojie and Wei, Xipin and Guo, Zhifang and Leng, Yichong and Lv, Yuanjun and He, Jinzheng and Lin, Junyang and Zhou, Chang and Zhou, Jingren},
journal={arXiv preprint arXiv:2407.10759},
year={2024}
}
@article{Qwen-Audio,
title={Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models},
author={Chu, Yunfei and Xu, Jin and Zhou, Xiaohuan and Yang, Qian and Zhang, Shiliang and Yan, Zhijie and Zhou, Chang and Zhou, Jingren},
journal={arXiv preprint arXiv:2311.07919},
year={2023}
}
Runs of Qwen Qwen2-Audio-7B on huggingface.co
36.2K
Total runs
-691
24-hour runs
-2.2K
3-day runs
-6.5K
7-day runs
-10.4K
30-day runs
More Information About Qwen2-Audio-7B huggingface.co Model
Qwen2-Audio-7B huggingface.co is an AI model on huggingface.co that provides Qwen2-Audio-7B's model effect (), which can be used instantly with this Qwen Qwen2-Audio-7B model. huggingface.co supports a free trial of the Qwen2-Audio-7B model, and also provides paid use of the Qwen2-Audio-7B. Support call Qwen2-Audio-7B model through api, including Node.js, Python, http.
Qwen2-Audio-7B huggingface.co is an online trial and call api platform, which integrates Qwen2-Audio-7B's modeling effects, including api services, and provides a free online trial of Qwen2-Audio-7B, you can try Qwen2-Audio-7B online for free by clicking the link below.
Qwen Qwen2-Audio-7B online free url in huggingface.co:
Qwen2-Audio-7B is an open source model from GitHub that offers a free installation service, and any user can find Qwen2-Audio-7B on GitHub to install. At the same time, huggingface.co provides the effect of Qwen2-Audio-7B install, users can directly use Qwen2-Audio-7B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.