LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
Introduction
LongCat-AudioDiT is a state-of-the-art (SOTA) diffusion-based text-to-speech (TTS) model that directly operates on the waveform latent space.
Abstract
: We present LongCat-TTS, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance.
Unlike previous methods that rely on intermediate acoustic representations such as mel-spectrograms, the core innovation of LongCat-TTS lies in operating directly within the waveform latent space. This approach effectively mitigates compounding errors and drastically simplifies the TTS pipeline, requiring only a waveform variational autoencoder (Wav-VAE) and a diffusion backbone.
Furthermore, we introduce two critical improvements to the inference process: first, we identify and rectify a long-standing training-inference mismatch; second, we replace traditional classifier-free guidance with adaptive projection guidance to elevate generation quality.
Experimental results demonstrate that, despite the absence of complex multi-stage training pipelines or high-quality human-annotated datasets, LongCat-TTS achieves SOTA zero-shot voice cloning performance on the Seed benchmark while maintaining competitive intelligibility.
Specifically, our largest variant, LongCat-TTS-3.5B, outperforms the previous SOTA model (Seed-TTS), improving the speaker similarity (SIM) scores from 0.809 to 0.818 on Seed-ZH, and from 0.776 to 0.797 on Seed-Hard.
Finally, through comprehensive ablation studies and systematic analysis, we validate the effectiveness of our proposed modules.
Notably, we investigate the interplay between the Wav-VAE and the TTS backbone, revealing the counterintuitive finding that superior reconstruction fidelity in the Wav-VAE does not necessarily lead to better overall TTS performance.
Code and model weights are released to foster further research within the speech community.
This repository provides the HuggingFace-compatible implementation, including model definition, weight conversion, and inference scripts.
Experimental Results on Seed Benchmark
LongCat-AudioDiT obtains state-of-the-art (SOTA) voice cloning performance on the Seed-benchmark, surpassing both close-source and open-source modles.
This repository, including both the model weights and the source code, is released under the
MIT License
.
Any contributions to this repository are licensed under the MIT License, unless otherwise stated. This license does not grant any rights to use Meituan trademarks or patents.
LongCat-AudioDiT-3.5B huggingface.co is an AI model on huggingface.co that provides LongCat-AudioDiT-3.5B's model effect (), which can be used instantly with this meituan-longcat LongCat-AudioDiT-3.5B model. huggingface.co supports a free trial of the LongCat-AudioDiT-3.5B model, and also provides paid use of the LongCat-AudioDiT-3.5B. Support call LongCat-AudioDiT-3.5B model through api, including Node.js, Python, http.
LongCat-AudioDiT-3.5B huggingface.co is an online trial and call api platform, which integrates LongCat-AudioDiT-3.5B's modeling effects, including api services, and provides a free online trial of LongCat-AudioDiT-3.5B, you can try LongCat-AudioDiT-3.5B online for free by clicking the link below.
meituan-longcat LongCat-AudioDiT-3.5B online free url in huggingface.co:
LongCat-AudioDiT-3.5B is an open source model from GitHub that offers a free installation service, and any user can find LongCat-AudioDiT-3.5B on GitHub to install. At the same time, huggingface.co provides the effect of LongCat-AudioDiT-3.5B install, users can directly use LongCat-AudioDiT-3.5B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
LongCat-AudioDiT-3.5B install url in huggingface.co: