from omnivoice import OmniVoice # code: github.com/k2-fsa/OmniVoice
model = OmniVoice.from_pretrained("multimodalart/omnivoice-word-control", device_map="cuda")
audio = model.generate(
text="I will <|pit_3|><|ton_ffall|>never agree to <|bnd_4|>this!",
language="en",
ref_audio="ref.wav",
ref_text="Transcript of the reference clip.",
)
The tokenizer in this repo already contains the 116 control tokens. For strict overall
timing, also pass
duration=
(the diffusion canvas is fixed before generation, so word
duration tokens redistribute time within it). Try it in the
demo Space
.
Training
Init from
k2-fsa/OmniVoice
; all parameters trained; masked-diffusion objective unchanged
(control tokens are conditioning only — no loss on text).
Data: WordVoice-5A
English
, all 4 train shards (~2,138 h, 1.18 M utterances); on-the-fly
tag construction with dropout (15% utterances untagged / 20% words / attributes kept @ 0.7)
to preserve plain TTS and enable partial control.
40k steps × 8,192 effective tokens (~1 epoch), lr 5e-5 cosine, bf16, single A100-40GB, ≈6 h.
Final eval loss 4.109 (best of run). Directional probes (hi/lo tag ratio, seed-matched):
pitch 3.61, energy 2.52, duration 2.70.
Fine-tuning recipe: the
word-control
patch set on the OmniVoice trainer
(
omnivoice/data/word_control.py
— tag vocabulary, binning, alignment, dropout).
License & lineage
Weights are
CC-BY-NC-4.0
(derivative of the CC-BY-NC OmniVoice release).
Data: WordVoice-5A (CC-BY-4.0). Task formulation: WordVoice, arXiv:2607.06461.
Runs of multimodalart omnivoice-word-control on huggingface.co
102
Total runs
0
24-hour runs
3
3-day runs
102
7-day runs
102
30-day runs
More Information About omnivoice-word-control huggingface.co Model
omnivoice-word-control huggingface.co is an AI model on huggingface.co that provides omnivoice-word-control's model effect (), which can be used instantly with this multimodalart omnivoice-word-control model. huggingface.co supports a free trial of the omnivoice-word-control model, and also provides paid use of the omnivoice-word-control. Support call omnivoice-word-control model through api, including Node.js, Python, http.
omnivoice-word-control huggingface.co is an online trial and call api platform, which integrates omnivoice-word-control's modeling effects, including api services, and provides a free online trial of omnivoice-word-control, you can try omnivoice-word-control online for free by clicking the link below.
multimodalart omnivoice-word-control online free url in huggingface.co:
omnivoice-word-control is an open source model from GitHub that offers a free installation service, and any user can find omnivoice-word-control on GitHub to install. At the same time, huggingface.co provides the effect of omnivoice-word-control install, users can directly use omnivoice-word-control installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
omnivoice-word-control install url in huggingface.co: