Note:
Piper uses the terms "medium", "high", etc. to refer to
model size
, not output quality.
Medium models (
63 MB, ~15M params) and high models (
114 MB, ~28M params) both produce 22.05 kHz audio.
Usage
With piper-tts (GPL)
from piper import PiperVoice
voice = PiperVoice.load("model.onnx")
for chunk in voice.synthesize("Hello, this is a test."):
# chunk.audio_float_array contains float32 audiopass
import json, subprocess, numpy as np, onnxruntime as ort, soundfile as sf
from huggingface_hub import hf_hub_download
model_id = "Trelis/piper-zh-cn-huayan-medium"
onnx_path = hf_hub_download(model_id, "model.onnx")
config_path = hf_hub_download(model_id, "model.onnx.json")
withopen(config_path) as f:
config = json.load(f)
session = ort.InferenceSession(onnx_path, providers=["CPUExecutionProvider"])
phoneme_id_map = config["phoneme_id_map"]
espeak_voice = config["espeak"]["voice"]
defphonemize(text, voice):
out = subprocess.run(
["espeak-ng", "-v", voice, "-q", "--ipa=2", "-x", text],
capture_output=True, text=True,
).stdout.strip()
return [list(line.replace("_", " ")) for line in out.split("\n") if line.strip()]
defto_ids(phonemes, pmap):
ids = [pmap["^"][0], pmap["_"][0]]
for p in phonemes:
if p in pmap:
ids.extend(pmap[p])
ids.append(pmap["_"][0])
ids.append(pmap["$"][0])
return ids
text = "Hello, this is a test."
audio_chunks = []
for sentence in phonemize(text, espeak_voice):
ids = to_ids(sentence, phoneme_id_map)
iflen(ids) < 3:
continue
audio = session.run(None, {
"input": np.array([ids], dtype=np.int64),
"input_lengths": np.array([len(ids)], dtype=np.int64),
"scales": np.array([
config["inference"]["noise_scale"],
config["inference"]["length_scale"],
config["inference"]["noise_w"],
], dtype=np.float32),
})[0]
audio_chunks.append(audio.squeeze())
audio = np.concatenate(audio_chunks).astype(np.float32)
sf.write("output.wav", audio, config["audio"]["sample_rate"])
Fine-tuning
You can fine-tune this model on your own voice data using
Trelis Studio
. Piper models can be trained on custom datasets to create personalized voices.
Attribution
Trained on data from
HuaYan TTS
. Fine-tuned from lessac medium.
piper-zh-cn-huayan-medium huggingface.co is an AI model on huggingface.co that provides piper-zh-cn-huayan-medium's model effect (), which can be used instantly with this Trelis piper-zh-cn-huayan-medium model. huggingface.co supports a free trial of the piper-zh-cn-huayan-medium model, and also provides paid use of the piper-zh-cn-huayan-medium. Support call piper-zh-cn-huayan-medium model through api, including Node.js, Python, http.
piper-zh-cn-huayan-medium huggingface.co is an online trial and call api platform, which integrates piper-zh-cn-huayan-medium's modeling effects, including api services, and provides a free online trial of piper-zh-cn-huayan-medium, you can try piper-zh-cn-huayan-medium online for free by clicking the link below.
Trelis piper-zh-cn-huayan-medium online free url in huggingface.co:
piper-zh-cn-huayan-medium is an open source model from GitHub that offers a free installation service, and any user can find piper-zh-cn-huayan-medium on GitHub to install. At the same time, huggingface.co provides the effect of piper-zh-cn-huayan-medium install, users can directly use piper-zh-cn-huayan-medium installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
piper-zh-cn-huayan-medium install url in huggingface.co: