espnet / fastspeech2_conformer

huggingface.co
Total runs: 6.8K
24-hour runs: 0
7-day runs: -577
30-day runs: -3.1K
Model's Last Updated: October 06 2023
text-to-audio

Introduction of fastspeech2_conformer

Model Details of fastspeech2_conformer

FastSpeech2Conformer

FastSpeech2Conformer is a non-autoregressive text-to-speech (TTS) model that combines the strengths of FastSpeech2 and the conformer architecture to generate high-quality speech from text quickly and efficiently.

Model Description

The FastSpeech2Conformer model was proposed with the paper Recent Developments On Espnet Toolkit Boosted By Conformer by Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, and Yuekai Zhang. It was first released in this repository . The license used is Apache 2.0 .

FastSpeech2 is a non-autoregressive TTS model, which means it can generate speech significantly faster than autoregressive models. It addresses some of the limitations of its predecessor, FastSpeech, by directly training the model with ground-truth targets instead of the simplified output from a teacher model. It also introduces more variation information of speech (e.g., pitch, energy, and more accurate duration) as conditional inputs. Furthermore, the conformer (convolutional transformer) architecture makes use of convolutions inside the transformer blocks to capture local speech patterns, while the attention layer is able to capture relationships in the input that are farther away.

  • Developed by: Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, and Yuekai Zhang.
  • Shared by: Connor Henderson
  • Model type: text-to-speech
  • Language(s) (NLP): [More Information Needed]
  • License: Apache 2.0
  • Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
🤗 Transformers Usage

You can run FastSpeech2Conformer locally with the 🤗 Transformers library.

  1. First install the 🤗 Transformers library , g2p-en:
pip install --upgrade pip
pip install --upgrade transformers g2p-en
  1. Run inference via the Transformers modelling code with the model and hifigan separately

from transformers import FastSpeech2ConformerTokenizer, FastSpeech2ConformerModel, FastSpeech2ConformerHifiGan
import soundfile as sf

tokenizer = FastSpeech2ConformerTokenizer.from_pretrained("espnet/fastspeech2_conformer")
inputs = tokenizer("Hello, my dog is cute.", return_tensors="pt")
input_ids = inputs["input_ids"]

model = FastSpeech2ConformerModel.from_pretrained("espnet/fastspeech2_conformer")
output_dict = model(input_ids, return_dict=True)
spectrogram = output_dict["spectrogram"]

hifigan = FastSpeech2ConformerHifiGan.from_pretrained("espnet/fastspeech2_conformer_hifigan")
waveform = hifigan(spectrogram)

sf.write("speech.wav", waveform.squeeze().detach().numpy(), samplerate=22050)
  1. Run inference via the Transformers modelling code with the model and hifigan combined
from transformers import FastSpeech2ConformerTokenizer, FastSpeech2ConformerWithHifiGan
import soundfile as sf

tokenizer = FastSpeech2ConformerTokenizer.from_pretrained("espnet/fastspeech2_conformer")
inputs = tokenizer("Hello, my dog is cute.", return_tensors="pt")
input_ids = inputs["input_ids"]

model = FastSpeech2ConformerWithHifiGan.from_pretrained("espnet/fastspeech2_conformer_with_hifigan")
output_dict = model(input_ids, return_dict=True)
waveform = output_dict["waveform"]

sf.write("speech.wav", waveform.squeeze().detach().numpy(), samplerate=22050)
  1. Run inference with a pipeline and specify which vocoder to use
from transformers import pipeline, FastSpeech2ConformerHifiGan
import soundfile as sf

vocoder = FastSpeech2ConformerHifiGan.from_pretrained("espnet/fastspeech2_conformer_hifigan")
synthesiser = pipeline(model="espnet/fastspeech2_conformer", vocoder=vocoder)

speech = synthesiser("Hello, my dog is cooler than you!")

sf.write("speech.wav", speech["audio"].squeeze(), samplerate=speech["sampling_rate"])
Direct Use

[More Information Needed]

Downstream Use [optional]

[More Information Needed]

Out-of-Scope Use

[More Information Needed]

Bias, Risks, and Limitations

[More Information Needed]

Recommendations

Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.

How to Get Started with the Model

Use the code below to get started with the model.

[More Information Needed]

Training Details
Training Data

[More Information Needed]

Training Procedure
Preprocessing [optional]

[More Information Needed]

Training Hyperparameters
  • Training regime: [More Information Needed]
Speeds, Sizes, Times [optional]

[More Information Needed]

Evaluation
Testing Data, Factors & Metrics
Testing Data

[More Information Needed]

Factors

[More Information Needed]

Metrics

[More Information Needed]

Results

[More Information Needed]

Summary
Model Examination [optional]

[More Information Needed]

Environmental Impact

Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019) .

  • Hardware Type: [More Information Needed]
  • Hours used: [More Information Needed]
  • Cloud Provider: [More Information Needed]
  • Compute Region: [More Information Needed]
  • Carbon Emitted: [More Information Needed]
Technical Specifications [optional]
Model Architecture and Objective

[More Information Needed]

Compute Infrastructure

[More Information Needed]

Hardware

[More Information Needed]

Software

[More Information Needed]

Citation [optional]

BibTeX:

[More Information Needed]

APA:

[More Information Needed]

Glossary [optional]

[More Information Needed]

More Information [optional]

[More Information Needed]

Model Card Authors [optional]

Connor Henderson (Disclaimer: no ESPnet affiliation)

Model Card Contact

[More Information Needed]

Runs of espnet fastspeech2_conformer on huggingface.co

6.8K
Total runs
0
24-hour runs
-149
3-day runs
-577
7-day runs
-3.1K
30-day runs

More Information About fastspeech2_conformer huggingface.co Model

More fastspeech2_conformer license Visit here:

https://choosealicense.com/licenses/apache-2.0

fastspeech2_conformer huggingface.co

fastspeech2_conformer huggingface.co is an AI model on huggingface.co that provides fastspeech2_conformer's model effect (), which can be used instantly with this espnet fastspeech2_conformer model. huggingface.co supports a free trial of the fastspeech2_conformer model, and also provides paid use of the fastspeech2_conformer. Support call fastspeech2_conformer model through api, including Node.js, Python, http.

fastspeech2_conformer huggingface.co Url

https://huggingface.co/espnet/fastspeech2_conformer

espnet fastspeech2_conformer online free

fastspeech2_conformer huggingface.co is an online trial and call api platform, which integrates fastspeech2_conformer's modeling effects, including api services, and provides a free online trial of fastspeech2_conformer, you can try fastspeech2_conformer online for free by clicking the link below.

espnet fastspeech2_conformer online free url in huggingface.co:

https://huggingface.co/espnet/fastspeech2_conformer

fastspeech2_conformer install

fastspeech2_conformer is an open source model from GitHub that offers a free installation service, and any user can find fastspeech2_conformer on GitHub to install. At the same time, huggingface.co provides the effect of fastspeech2_conformer install, users can directly use fastspeech2_conformer installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

fastspeech2_conformer install url in huggingface.co:

https://huggingface.co/espnet/fastspeech2_conformer

Url of fastspeech2_conformer

fastspeech2_conformer huggingface.co Url

Provider of fastspeech2_conformer huggingface.co

espnet
ORGANIZATIONS

Other API from espnet

huggingface.co

Total runs: 12.1K
Run Growth: 6.9K
Growth Rate: 56.27%
Updated:September 20 2026
huggingface.co

Total runs: 5.5K
Run Growth: 3.7K
Growth Rate: 64.87%
Updated:September 18 2026
huggingface.co

Total runs: 183
Run Growth: 124
Growth Rate: 67.76%
Updated:January 22 2026
huggingface.co

Total runs: 116
Run Growth: 64
Growth Rate: 54.70%
Updated:September 20 2026
huggingface.co

Total runs: 73
Run Growth: 6
Growth Rate: 8.45%
Updated:June 17 2025
huggingface.co

Total runs: 67
Run Growth: -12
Growth Rate: -17.91%
Updated:September 20 2026
huggingface.co

Total runs: 21
Run Growth: 15
Growth Rate: 71.43%
Updated:September 20 2026
huggingface.co

Total runs: 18
Run Growth: 16
Growth Rate: 88.89%
Updated:September 20 2026