The model generates speech tokens autoregressively — the LM produces
<|vision_pad|>
(speech_diffusion) tokens that trigger diffusion sampling,
with
<|vision_start|>
/
<|vision_end|>
as control tokens.
Quality
Input
Parakeet ASR
"Hello, how are you today?"
"Hello, how are you today?"
Differences from Realtime-0.5B
Feature
Realtime-0.5B
1.5B Base
Architecture
4L base + 20L TTS LM
Single 28L LM
Voice input
Pre-computed .pt prompts
Audio WAV files
Voice cloning
No (fixed presets)
Yes (from reference audio)
Multi-speaker
No
Yes (up to 4 speakers)
Streaming
Yes
No
License
MIT (same as original model).
Runs of cstr vibevoice-1.5b-GGUF on huggingface.co
1.1K
Total runs
40
24-hour runs
70
3-day runs
-949
7-day runs
-1.5K
30-day runs
More Information About vibevoice-1.5b-GGUF huggingface.co Model
vibevoice-1.5b-GGUF huggingface.co is an AI model on huggingface.co that provides vibevoice-1.5b-GGUF's model effect (), which can be used instantly with this cstr vibevoice-1.5b-GGUF model. huggingface.co supports a free trial of the vibevoice-1.5b-GGUF model, and also provides paid use of the vibevoice-1.5b-GGUF. Support call vibevoice-1.5b-GGUF model through api, including Node.js, Python, http.
vibevoice-1.5b-GGUF huggingface.co is an online trial and call api platform, which integrates vibevoice-1.5b-GGUF's modeling effects, including api services, and provides a free online trial of vibevoice-1.5b-GGUF, you can try vibevoice-1.5b-GGUF online for free by clicking the link below.
cstr vibevoice-1.5b-GGUF online free url in huggingface.co:
vibevoice-1.5b-GGUF is an open source model from GitHub that offers a free installation service, and any user can find vibevoice-1.5b-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of vibevoice-1.5b-GGUF install, users can directly use vibevoice-1.5b-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
vibevoice-1.5b-GGUF install url in huggingface.co: