k2 is necessary for training
and can speed up inference. Nevertheless, you can still use the inference mode of ZipVoice without installing k2.
Note:
Make sure to install the k2 version that matches your PyTorch and CUDA version. For example, if you are using pytorch 2.5.1 and CUDA 12.1, you can install k2 as follows:
To generate single-speaker speech with our pre-trained ZipVoice or ZipVoice-Distill models, use the following commands (Required models will be downloaded from HuggingFace):
1.1 Inference of a single sentence
python3 -m zipvoice.bin.infer_zipvoice \
--model-name zipvoice \
--prompt-wav prompt.wav \
--prompt-text "I am the transcription of the prompt wav." \
--text "I am the text to be synthesized." \
--res-wav-path result.wav
--model-name
can be
zipvoice
or
zipvoice_distill
, which are models before and after distillation, respectively.
If
<>
or
[]
appear in the text, strings enclosed by them will be treated as special tokens.
<>
denotes Chinese pinyin and
[]
denotes other special tags.
Could run ONNX models on CPU faster with
zipvoice.bin.infer_zipvoice_onnx
.
Note:
If you have trouble connecting to HuggingFace, try:
Each line of
test.tsv
is in the format of
{wav_name}\t{prompt_transcription}\t{prompt_wav}\t{text}
.
2. Dialogue speech generation
2.1 Inference command
To generate two-party spoken dialogues with our pre-trained ZipVoice-Dialogue or ZipVoice-Dialogue-Stereo models, use the following commands (Required models will be downloaded from HuggingFace):
You can also scan the QR code to join our wechat group or follow our wechat official account.
Wechat Group
Wechat Official Account
Citation
@article{zhu2025zipvoice,
title={ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching},
author={Zhu, Han and Kang, Wei and Yao, Zengwei and Guo, Liyong and Kuang, Fangjun and Li, Zhaoqing and Zhuang, Weiji and Lin, Long and Povey, Daniel},
journal={arXiv preprint arXiv:2506.13053},
year={2025}
}
@article{zhu2025zipvoicedialog,
title={ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching},
author={Zhu, Han and Kang, Wei and Guo, Liyong and Yao, Zengwei and Kuang, Fangjun and Zhuang, Weiji and Li, Zhaoqing and Han, Zhifeng and Zhang, Dong and Zhang, Xin and Song, Xingchen and Lin, Long and Povey, Daniel},
journal={arXiv preprint arXiv:2507.09318},
year={2025}
}
Runs of Respair Zip_dial on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About Zip_dial huggingface.co Model
Zip_dial huggingface.co
Zip_dial huggingface.co is an AI model on huggingface.co that provides Zip_dial's model effect (), which can be used instantly with this Respair Zip_dial model. huggingface.co supports a free trial of the Zip_dial model, and also provides paid use of the Zip_dial. Support call Zip_dial model through api, including Node.js, Python, http.
Zip_dial huggingface.co is an online trial and call api platform, which integrates Zip_dial's modeling effects, including api services, and provides a free online trial of Zip_dial, you can try Zip_dial online for free by clicking the link below.
Respair Zip_dial online free url in huggingface.co:
Zip_dial is an open source model from GitHub that offers a free installation service, and any user can find Zip_dial on GitHub to install. At the same time, huggingface.co provides the effect of Zip_dial install, users can directly use Zip_dial installed effect in huggingface.co for debugging and trial. It also supports api for free installation.