VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
VITS is an end-to-end speech synthesis model that predicts a speech waveform conditional on an input text sequence. It is a
conditional variational autoencoder (VAE) comprised of a posterior encoder, decoder, and conditional prior. This repository
contains the weights for the official VITS checkpoint trained on the
LJ Speech
dataset.
Model Details
VITS (
V
ariational
I
nference with adversarial learning for end-to-end
T
ext-to-
S
peech) is an end-to-end
speech synthesis model that predicts a speech waveform conditional on an input text sequence. It is a conditional variational
autoencoder (VAE) comprised of a posterior encoder, decoder, and conditional prior.
A set of spectrogram-based acoustic features are predicted by the flow-based module, which is formed of a Transformer-based
text encoder and multiple coupling layers. The spectrogram is decoded using a stack of transposed convolutional layers,
much in the same style as the HiFi-GAN vocoder. Motivated by the one-to-many nature of the TTS problem, where the same text
input can be spoken in multiple ways, the model also includes a stochastic duration predictor, which allows the model to
synthesise speech with different rhythms from the same input text.
The model is trained end-to-end with a combination of losses derived from variational lower bound and adversarial training.
To improve the expressiveness of the model, normalizing flows are applied to the conditional prior distribution. During
inference, the text encodings are up-sampled based on the duration prediction module, and then mapped into the
waveform using a cascade of the flow module and HiFi-GAN decoder. Due to the stochastic nature of the duration predictor,
the model is non-deterministic, and thus requires a fixed seed to generate the same speech waveform.
There are two variants of the VITS model: one is trained on the
LJ Speech
dataset,
and the other is trained on the
VCTK
dataset. LJ Speech dataset consists of 13,100 short
audio clips of a single speaker with a total length of approximately 24 hours. The VCTK dataset consists of approximately 44,000
short audio clips uttered by 109 native English speakers with various accents. The total length of the audio clips is approximately
44 hours.
VITS is available in the 🤗 Transformers library from version 4.33 onwards. To use this checkpoint,
first install the latest version of the library:
pip install --upgrade transformers accelerate
Then, run inference with the following code-snippet:
from transformers import VitsModel, AutoTokenizer
import torch
model = VitsModel.from_pretrained("kakao-enterprise/vits-ljs")
tokenizer = AutoTokenizer.from_pretrained("kakao-enterprise/vits-ljs")
text = "Hey, it's Hugging Face on the phone"
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
output = model(**inputs).waveform
The resulting waveform can be saved as a
.wav
file:
vits-ljs huggingface.co is an AI model on huggingface.co that provides vits-ljs's model effect (), which can be used instantly with this kakao-enterprise vits-ljs model. huggingface.co supports a free trial of the vits-ljs model, and also provides paid use of the vits-ljs. Support call vits-ljs model through api, including Node.js, Python, http.
vits-ljs huggingface.co is an online trial and call api platform, which integrates vits-ljs's modeling effects, including api services, and provides a free online trial of vits-ljs, you can try vits-ljs online for free by clicking the link below.
kakao-enterprise vits-ljs online free url in huggingface.co:
vits-ljs is an open source model from GitHub that offers a free installation service, and any user can find vits-ljs on GitHub to install. At the same time, huggingface.co provides the effect of vits-ljs install, users can directly use vits-ljs installed effect in huggingface.co for debugging and trial. It also supports api for free installation.