🐁POWSM is the first phonetic foundation model that can perform four phone-related tasks:
Phone Recognition (PR), Automatic Speech Recognition (ASR), audio-guided grapheme-to-phoneme conversion (G2P), and audio-guided phoneme-to-grapheme
conversion (P2G).
Our models are trained on 16kHz audio with a fixed duration of 20s. When using the pre-trained model, please ensure the input speech is 16kHz and pad or truncate it to 20s.
To distinguish phone entries from BPE tokens that share the same Unicode, we enclose every phone in slashes and treat them as special tokens. For example, /pʰɔsəm/ would be tokenized as /pʰ//ɔ//s//ə//m/.
from espnet2.bin.s2t_inference import Speech2Text
import soundfile as sf # or librosa
task = '<pr>'
s2t = Speech2Text.from_pretrained(
"espnet/powsm",
device="cuda",
lang_sym='<eng>', # ISO 639-3; set to <unk> for unseen languages
task_sym=task, # <pr>, <asr>, <g2p>, <p2g>
)
speech, rate = sf.read("sample.wav", sr=16000)
prompt = "<na>"# G2P: set to ASR transcript; P2G: set to phone transcription with slashes
pred = s2t(speech, text_prev=prompt)[0][0]
if task == '<pr>'or task == '<g2p>: pred = pred.replace("/", "")print(pred)
Other tasks
See
force_align.py
in
ESPnet recipe
to try out CTC forced alignment with POWSM's encoder!
LID is learned implicitly during training, and you may run it with the script below:
from espnet2.bin.s2t_inference_language import Speech2Language
import soundfile as sf # or librosa
s2t = Speech2Language.from_pretrained(
"espnet/powsm",
device="cuda",
nbest=1, # number of possible languages to return
first_lang_sym="<afr>", # fixed; defined in vocab list
last_lang_sym="<zul>"# fixed; defined in vocab list
)
speech, rate = sf.read("sample.wav", sr=16000)
pred = model(speech)[0] # a list of lang-prob pairprint(pred)
powsm huggingface.co is an AI model on huggingface.co that provides powsm's model effect (), which can be used instantly with this espnet powsm model. huggingface.co supports a free trial of the powsm model, and also provides paid use of the powsm. Support call powsm model through api, including Node.js, Python, http.
powsm huggingface.co is an online trial and call api platform, which integrates powsm's modeling effects, including api services, and provides a free online trial of powsm, you can try powsm online for free by clicking the link below.
powsm is an open source model from GitHub that offers a free installation service, and any user can find powsm on GitHub to install. At the same time, huggingface.co provides the effect of powsm install, users can directly use powsm installed effect in huggingface.co for debugging and trial. It also supports api for free installation.