MeanVC2
is a robust, low-latency streaming zero-shot voice conversion (VC) system built upon the diffusion-based conditional flow matching (CFM) framework. By introducing
Future-Receptive Chunking (FRC)
and a
Universal Timbre Token Encoder (UTTE)
, MeanVC2 achieves high-fidelity voice conversion with an end-to-end pipeline latency of only
110 ms
while maintaining superior speaker similarity and audio naturalness even with a
40 ms chunk size
.
✨ Key Features
🚀 Ultra-Low Latency Streaming
: 110 ms end-to-end first-packet latency with 40 ms chunk size; full pipeline RTF < 0.633 on single CPU core.
⚡ Single-Step Generation
: Mean flows + 1-NFE ODE solving for high-quality mel-spectrogram synthesis.
🎯 Zero-Shot Capability
: Convert to any unseen target speaker without re-training, robust under low-quality reference audio.
💾 Lightweight
: Only 18M parameters — far smaller than competing streaming VC systems.
🔊 High Fidelity
: Future-Receptive Chunking (FRC) with Universal Timbre Token Encoder (UTTE) for pronunciation-aware timbre modeling.
📁 Model Checkpoints
This repository hosts all pretrained models for MeanVC2:
Voice Conversion Models
File
Description
meanvc2_120ms_40ms.safetensors
120ms chunk + 40ms future (recommended for quality)
meanvc2_40ms_40ms.safetensors
40ms chunk + 40ms future (lower latency)
Vocoder (JIT)
File
Description
vocos.pt
Vocos JIT-traced vocoder
ASR Encoder (Fast-U2++ JIT)
File
Description
fastu2pp_80ms.pt
80ms chunk JIT model (11-frame window, stride=8)
fastu2pp_160ms.pt
160ms chunk JIT model (19-frame window, stride=16)
Fast-U2++ (WeNet) extracts bottleneck features (BNFs) from source waveform
Speaker Encoder
ECAPA-TDNN + WavLM upstream extracts global speaker embedding from reference audio
Universal Timbre Token Encoder (UTTE)
Transforms speaker embedding into K key-value UTT pairs; BNFs serve as queries in cross-attention for fine-grained, pronunciation-aware timbre cues
DiT-based CFM Decoder
4-layer DiT (hidden dim 512, 2 heads) with Future-Receptive Chunking (FRC); trained with mean flows objective for 1-NFE mel-spectrogram generation
Vocoder
Vocos converts mel-spectrograms to 16 kHz high-fidelity speech waveforms
Total parameters
: ~18M
📜 License & Disclaimer
MeanVC2 is released under the
Apache License 2.0
. This open-source license allows you to freely use, modify, and distribute the model, as long as you include the appropriate copyright notice and disclaimer.
MeanVC2 is designed for research and legitimate applications in voice conversion technology. Users must obtain proper consent from individuals whose voices are being converted or used as references. We strongly discourage malicious use including impersonation, fraud, or creating misleading audio content. Users are solely responsible for ensuring compliance with ethical standards and legal requirements.
📄 Citation
If you find our work helpful, please cite:
@article{ma2026meanvc2,
title={MeanVC2: Robust Low-Latency Streaming Zero-Shot Voice Conversion},
author={Ma, Guobin and Xia, Yuxuan and Jiang, Yuepeng and Guo, Dake and Xie, Hanke and Hu, Jingbin and Wang, Yanbo and Xie, Lei and Zhu, Pengcheng},
journal={arXiv preprint arXiv:2606.09050},
year={2026}
}
MeanVC2 huggingface.co is an AI model on huggingface.co that provides MeanVC2's model effect (), which can be used instantly with this ASLP-lab MeanVC2 model. huggingface.co supports a free trial of the MeanVC2 model, and also provides paid use of the MeanVC2. Support call MeanVC2 model through api, including Node.js, Python, http.
MeanVC2 huggingface.co is an online trial and call api platform, which integrates MeanVC2's modeling effects, including api services, and provides a free online trial of MeanVC2, you can try MeanVC2 online for free by clicking the link below.
ASLP-lab MeanVC2 online free url in huggingface.co:
MeanVC2 is an open source model from GitHub that offers a free installation service, and any user can find MeanVC2 on GitHub to install. At the same time, huggingface.co provides the effect of MeanVC2 install, users can directly use MeanVC2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.