GLM-TTS is a high-quality text-to-speech (TTS) synthesis system based on large language models, supporting zero-shot voice cloning and streaming inference. The system adopts a two-stage architecture combining an LLM for speech token generation and a Flow Matching model for waveform synthesis.
By introducing a
Multi-Reward Reinforcement Learning
framework, GLM-TTS significantly improves the expressiveness of generated speech, achieving more natural emotional control compared to traditional TTS systems.
Key Features
Zero-shot Voice Cloning:
Clone any speaker's voice with just 3-10 seconds of prompt audio.
RL-enhanced Emotion Control:
Utilizes a multi-reward reinforcement learning framework (GRPO) to optimize prosody and emotion.
High-quality Synthesis:
Generates speech comparable to commercial systems with reduced Character Error Rate (CER).
Phoneme-level Control:
Supports "Hybrid Phoneme + Text" input for precise pronunciation control (e.g., polyphones).
Streaming Inference:
Supports real-time audio generation suitable for interactive applications.
Bilingual Support:
Optimized for Chinese and English mixed text.
System Architecture
GLM-TTS follows a two-stage design:
Stage 1 (LLM):
A Llama-based model converts input text into speech token sequences.
Stage 2 (Flow Matching):
A Flow model converts token sequences into high-quality mel-spectrograms, which are then turned into waveforms by a vocoder.
Reinforcement Learning Alignment
To tackle flat emotional expression, GLM-TTS uses a
Group Relative Policy Optimization (GRPO)
algorithm with multiple reward functions (Similarity, CER, Emotion, Laughter) to align the LLM's generation strategy.
Evaluation Results
Evaluated on
seed-tts-eval
.
GLM-TTS_RL
achieves the lowest Character Error Rate (CER) while maintaining high speaker similarity.
Model
CER ↓
SIM ↑
Open-source
Seed-TTS
1.12
79.6
🔒 No
CosyVoice2
1.38
75.7
👐 Yes
F5-TTS
1.53
76.0
👐 Yes
GLM-TTS (Base)
1.03
76.1
👐 Yes
GLM-TTS_RL (Ours)
0.89
76.4
👐 Yes
Quick Start
Installation
git clone [https://github.com/zai-org/GLM-TTS.git](https://github.com/zai-org/GLM-TTS.git)
cd GLM-TTS
pip install -r requirements.txt
Command Line Inference
python glmtts_inference.py \
--data=example_zh \
--exp_name=_test \
--use_cache \
# --phoneme # Add this flag to enable phoneme capabilities.
Shell Script Inference
bash glmtts_inference.sh
Acknowledgments & Citation
We thank the following open-source projects for their support:
CosyVoice
- Providing frontend processing framework and high-quality vocoder
Llama
- Providing basic language model architecture
GLM-TTS huggingface.co is an AI model on huggingface.co that provides GLM-TTS's model effect (), which can be used instantly with this zai-org GLM-TTS model. huggingface.co supports a free trial of the GLM-TTS model, and also provides paid use of the GLM-TTS. Support call GLM-TTS model through api, including Node.js, Python, http.
GLM-TTS huggingface.co is an online trial and call api platform, which integrates GLM-TTS's modeling effects, including api services, and provides a free online trial of GLM-TTS, you can try GLM-TTS online for free by clicking the link below.
zai-org GLM-TTS online free url in huggingface.co:
GLM-TTS is an open source model from GitHub that offers a free installation service, and any user can find GLM-TTS on GitHub to install. At the same time, huggingface.co provides the effect of GLM-TTS install, users can directly use GLM-TTS installed effect in huggingface.co for debugging and trial. It also supports api for free installation.