VibeVoice-ASR: Long-Form Rich Transcription with User Prompts
VibeVoice-ASR
is the latest addition to the
VibeVoice
family. While the original VibeVoice / VibeVoice-Realtime focused on expressive TTS,
VibeVoice-ASR
focuses on understanding long-form speech with high precision and rich metadata.
It is a unified speech-to-text model designed to handle
1-hour long-form audio
in a single pass, generating structured transcriptions containing
Who (Speaker), When (Timestamps), and What (Content)
, with support for
User-Customized Context
.
🕒 60-min Single-Pass Processing
:
Unlike conventional ASR models that slice audio into short chunks (often losing global context), VibeVoice ASR accepts up to
60 minutes
of continuous audio input within 64K length. This ensures consistent speaker tracking and semantic coherence across the entire hour.
👤 Optional Context Injection
:
Users can provide customized context (e.g., specific names, technical terms, or background info) to guide the recognition process, significantly improving accuracy on domain-specific content.
📝 Rich Transcription (Who, When, What)
:
The model performs ASR, Diarization, and Timestamping simultaneously. The output is a structured sequence indicating
who
said
what
at
which time
.
This project was conducted by members of Microsoft Research. We welcome feedback and collaboration from our audience. If you have suggestions, questions, or observe unexpected/offensive behavior in our technology, please contact us at
[email protected]
.
If the team receives reports of undesired behavior or identifies issues independently, we will update this repository with appropriate mitigations.
Runs of microsoft VibeVoice-ASR on huggingface.co
721.4K
Total runs
1.4K
24-hour runs
3.6K
3-day runs
16.7K
7-day runs
24.8K
30-day runs
More Information About VibeVoice-ASR huggingface.co Model
VibeVoice-ASR huggingface.co is an AI model on huggingface.co that provides VibeVoice-ASR's model effect (), which can be used instantly with this microsoft VibeVoice-ASR model. huggingface.co supports a free trial of the VibeVoice-ASR model, and also provides paid use of the VibeVoice-ASR. Support call VibeVoice-ASR model through api, including Node.js, Python, http.
VibeVoice-ASR huggingface.co is an online trial and call api platform, which integrates VibeVoice-ASR's modeling effects, including api services, and provides a free online trial of VibeVoice-ASR, you can try VibeVoice-ASR online for free by clicking the link below.
microsoft VibeVoice-ASR online free url in huggingface.co:
VibeVoice-ASR is an open source model from GitHub that offers a free installation service, and any user can find VibeVoice-ASR on GitHub to install. At the same time, huggingface.co provides the effect of VibeVoice-ASR install, users can directly use VibeVoice-ASR installed effect in huggingface.co for debugging and trial. It also supports api for free installation.