VoxCPM2
is a tokenizer-free, diffusion autoregressive Text-to-Speech model —
2B parameters
,
30 languages
,
48kHz
audio output, trained on over
2 million hours
of multilingual speech data.
Highlights
🌍
30-Language Multilingual
— No language tag needed; input text in any supported language directly
🎨
Voice Design
— Generate a novel voice from a natural-language description alone (gender, age, tone, emotion, pace…); no reference audio required
🎛️
Controllable Cloning
— Clone any voice from a short clip, with optional style guidance to steer emotion, pace, and expression while preserving timbre
🎙️
Ultimate Cloning
— Provide reference audio + its transcript for audio-continuation cloning; every vocal nuance faithfully reproduced
🔊
48kHz Studio-Quality Output
— Accepts 16kHz reference; outputs 48kHz via AudioVAE V2's built-in super-resolution, no external upsampler needed
🧠
Context-Aware Synthesis
— Automatically infers appropriate prosody and expressiveness from text content
⚡
Real-Time Streaming
— RTF as low as ~0.3 on NVIDIA RTX 4090, and ~0.13 accelerated by
Nano-VLLM
📜
Fully Open-Source & Commercial-Ready
— Apache-2.0 license, free for commercial use
from voxcpm import VoxCPM
import soundfile as sf
model = VoxCPM.from_pretrained("openbmb/VoxCPM2", load_denoiser=False)
wav = model.generate(
text="VoxCPM2 brings multilingual support, creative voice design, and controllable voice cloning.",
cfg_value=2.0,
inference_timesteps=10,
)
sf.write("output.wav", wav, model.tts_model.sample_rate)
Voice Design
Put the voice description in parentheses at the start of
text
, followed by the content to synthesize:
wav = model.generate(
text="(A young woman, gentle and sweet voice)Hello, welcome to VoxCPM2!",
cfg_value=2.0,
inference_timesteps=10,
)
sf.write("voice_design.wav", wav, model.tts_model.sample_rate)
Controllable Voice Cloning
# Basic cloning
wav = model.generate(
text="This is a cloned voice generated by VoxCPM2.",
reference_wav_path="speaker.wav",
)
sf.write("clone.wav", wav, model.tts_model.sample_rate)
# Cloning with style control
wav = model.generate(
text="(slightly faster, cheerful tone)This is a cloned voice with style control.",
reference_wav_path="speaker.wav",
cfg_value=2.0,
inference_timesteps=10,
)
sf.write("controllable_clone.wav", wav, model.tts_model.sample_rate)
Ultimate Cloning
Provide both the reference audio and its exact transcript for maximum fidelity. Pass the same clip to both
reference_wav_path
and
prompt_wav_path
for highest similarity:
wav = model.generate(
text="This is an ultimate cloning demonstration using VoxCPM2.",
prompt_wav_path="speaker_reference.wav",
prompt_text="The transcript of the reference audio.",
reference_wav_path="speaker_reference.wav",
)
sf.write("hifi_clone.wav", wav, model.tts_model.sample_rate)
Streaming
import numpy as np
chunks = []
for chunk in model.generate_streaming(text="Streaming is easy with VoxCPM!"):
chunks.append(chunk)
wav = np.concatenate(chunks)
sf.write("streaming.wav", wav, model.tts_model.sample_rate)
Voice Design and Style Control results may vary between runs; generating 1–3 times is recommended to obtain the desired output.
Performance varies across languages depending on training data availability.
Occasional instability may occur with very long or highly expressive inputs.
Strictly forbidden
to use for impersonation, fraud, or disinformation. AI-generated content should be clearly labeled.
Citation
@article{voxcpm2_2026,
title = {VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning},
author = {VoxCPM Team},
journal = {GitHub},
year = {2026},
}
@article{voxcpm2025,
title = {VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning},
author = {Zhou, Yixuan and Zeng, Guoyang and Liu, Xin and Li, Xiang and
Yu, Renjie and Wang, Ziyang and Ye, Runchuan and Sun, Weiyue and
Gui, Jiancheng and Li, Kehan and Wu, Zhiyong and Liu, Zhiyuan},
journal = {arXiv preprint arXiv:2509.24650},
year = {2025},
}
License
Released under the
Apache-2.0
license, free for commercial use. For production deployments, we recommend thorough testing and safety evaluation tailored to your use case.
Runs of eadx VoxCPM2 on huggingface.co
12
Total runs
0
24-hour runs
1
3-day runs
3
7-day runs
11
30-day runs
More Information About VoxCPM2 huggingface.co Model
VoxCPM2 huggingface.co is an AI model on huggingface.co that provides VoxCPM2's model effect (), which can be used instantly with this eadx VoxCPM2 model. huggingface.co supports a free trial of the VoxCPM2 model, and also provides paid use of the VoxCPM2. Support call VoxCPM2 model through api, including Node.js, Python, http.
VoxCPM2 huggingface.co is an online trial and call api platform, which integrates VoxCPM2's modeling effects, including api services, and provides a free online trial of VoxCPM2, you can try VoxCPM2 online for free by clicking the link below.
VoxCPM2 is an open source model from GitHub that offers a free installation service, and any user can find VoxCPM2 on GitHub to install. At the same time, huggingface.co provides the effect of VoxCPM2 install, users can directly use VoxCPM2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.