microsoft / VibeVoice-AcousticTokenizer

huggingface.co
Total runs: 529
24-hour runs: 2
7-day runs: 16
30-day runs: -423
Model's Last Updated: February 06 2026
feature-extraction

Introduction of VibeVoice-AcousticTokenizer

Model Details of VibeVoice-AcousticTokenizer

VibeVoice Acoustic Tokenizer

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.

A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a Large Language Model (LLM) to understand textual context and dialogue flow, and a diffusion head to generate high-fidelity acoustic details.

The speech tokenizer is a key component for both VibeVoice TTS and ASR .

➡️ Technical Report: VibeVoice Technical Report

➡️ Project Page: microsoft/VibeVoice

Tokenizer Comparison

Models

Model Context Length Length (min) Weight
VibeVoice-Realtime-0.5B 8K ~10 min HF link
VibeVoice-1.5B 64K ~90 min HF link
VibeVoice-ASR 64K ~60 min HF link
VibeVoice-AcousticTokenizer - - This model

Usage

Setup

Until the VibeVoice acoustic tokenizer is part of an official Transformers release, it can be used by installing from the source code:

pip install git+https://github.com/huggingface/transformers.git
Example
Encoding and decoding
import torch
from scipy.io import wavfile

from transformers import AutoFeatureExtractor, VibeVoiceAcousticTokenizerModel
from transformers.audio_utils import load_audio_librosa


model_id = "microsoft/VibeVoice-AcousticTokenizer"

# load model
feature_extractor = AutoFeatureExtractor.from_pretrained(model_id)
model = VibeVoiceAcousticTokenizerModel.from_pretrained(model_id, device_map="auto")
print("Model loaded on device:", model.device)
print("Model dtype:", model.dtype)

# load audio
audio = load_audio_librosa(
    "https://huggingface.co/datasets/bezzam/vibevoice_samples/resolve/main/voices/en-Alice_woman.wav",
    sampling_rate=feature_extractor.sampling_rate,
)

# preprocess audio
inputs = feature_extractor(
    audio,
    sampling_rate=feature_extractor.sampling_rate,
    pad_to_multiple_of=3200,
).to(model.device, model.dtype)
print("Input audio shape:", inputs.input_values.shape)
# Input audio shape: torch.Size([1, 1, 224000])

with torch.no_grad():
    # set VAE sampling to False for deterministic output
    encoded_outputs = model.encode(inputs.input_values, sample=False)
    print("Latent shape:", encoded_outputs.latents.shape)
    # Latent shape: torch.Size([1, 70, 64])

    decoded_outputs = model.decode(**encoded_outputs)
    print("Reconstructed audio shape:", decoded_outputs.audio.shape)
    # Reconstructed audio shape: torch.Size([1, 1, 224000])

# Save audio
output_fp = "vibevoice_acoustic_tokenizer_reconstructed.wav"
wavfile.write(output_fp, feature_extractor.sampling_rate, decoded_outputs.audio.squeeze().float().cpu().numpy())
print(f"Reconstructed audio saved to : {output_fp}")

Original audio

Encoded/decoded audio

Streaming

For streaming ASR or TTS, where cached states need to be tracked, the use_cache parameter can be used when encoding or decoding audio:

import torch
from scipy.io import wavfile

from transformers import AutoFeatureExtractor, VibeVoiceAcousticTokenizerModel
from transformers.audio_utils import load_audio_librosa


model_id = "microsoft/VibeVoice-AcousticTokenizer"


# load model
feature_extractor = AutoFeatureExtractor.from_pretrained(model_id)
model = VibeVoiceAcousticTokenizerModel.from_pretrained(model_id, device_map="auto")
print("Model loaded on device:", model.device)
print("Model dtype:", model.dtype)

# load audio
audio = load_audio_librosa(
    "https://huggingface.co/datasets/bezzam/vibevoice_samples/resolve/main/voices/en-Alice_woman.wav",
    sampling_rate=feature_extractor.sampling_rate,
)

# preprocess audio
inputs = feature_extractor(
    audio,
    sampling_rate=feature_extractor.sampling_rate,
    pad_to_multiple_of=3200,
).to(model.device, model.dtype)
print("Input audio shape:", inputs.input_values.shape)
# Input audio shape: torch.Size([1, 1, 224000])

# chache will be initialized after a first pass
encoder_cache = None
decoder_cache = None
with torch.no_grad():
    # set VAE sampling to False for deterministic output
    encoded_outputs = model.encode(inputs.input_values, sample=False, padding_cache=encoder_cache, use_cache=True)
    print("Latent shape:", encoded_outputs.latents.shape)
    # Latent shape: torch.Size([1, 70, 64])

    decoded_outputs = model.decode(encoded_outputs.latents, padding_cache=decoder_cache, use_cache=True)
    print("Reconstructed audio shape:", decoded_outputs.audio.shape)
    # Reconstructed audio shape: torch.Size([1, 1, 224000])

    # `padding_cache` can be extracted from the outputs for subsequent passes
    encoder_cache = encoded_outputs.padding_cache
    print("Number of cached encoder layers:", len(encoder_cache.per_layer_in_channels))
    # Number of cached encoder layers: 34
    decoder_cache = decoded_outputs.padding_cache
    print("Number of cached decoder layers:", len(decoder_cache.per_layer_in_channels))
    # Number of cached decoder layers: 34

# Save audio
output_fp = "vibevoice_acoustic_tokenizer_reconstructed.wav"
wavfile.write(output_fp, feature_extractor.sampling_rate, decoded_outputs.audio.squeeze().float().cpu().numpy())
print(f"Reconstructed audio saved to : {output_fp}")

Runs of microsoft VibeVoice-AcousticTokenizer on huggingface.co

529
Total runs
2
24-hour runs
13
3-day runs
16
7-day runs
-423
30-day runs

More Information About VibeVoice-AcousticTokenizer huggingface.co Model

More VibeVoice-AcousticTokenizer license Visit here:

https://choosealicense.com/licenses/mit

VibeVoice-AcousticTokenizer huggingface.co

VibeVoice-AcousticTokenizer huggingface.co is an AI model on huggingface.co that provides VibeVoice-AcousticTokenizer's model effect (), which can be used instantly with this microsoft VibeVoice-AcousticTokenizer model. huggingface.co supports a free trial of the VibeVoice-AcousticTokenizer model, and also provides paid use of the VibeVoice-AcousticTokenizer. Support call VibeVoice-AcousticTokenizer model through api, including Node.js, Python, http.

VibeVoice-AcousticTokenizer huggingface.co Url

https://huggingface.co/microsoft/VibeVoice-AcousticTokenizer

microsoft VibeVoice-AcousticTokenizer online free

VibeVoice-AcousticTokenizer huggingface.co is an online trial and call api platform, which integrates VibeVoice-AcousticTokenizer's modeling effects, including api services, and provides a free online trial of VibeVoice-AcousticTokenizer, you can try VibeVoice-AcousticTokenizer online for free by clicking the link below.

microsoft VibeVoice-AcousticTokenizer online free url in huggingface.co:

https://huggingface.co/microsoft/VibeVoice-AcousticTokenizer

VibeVoice-AcousticTokenizer install

VibeVoice-AcousticTokenizer is an open source model from GitHub that offers a free installation service, and any user can find VibeVoice-AcousticTokenizer on GitHub to install. At the same time, huggingface.co provides the effect of VibeVoice-AcousticTokenizer install, users can directly use VibeVoice-AcousticTokenizer installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

VibeVoice-AcousticTokenizer install url in huggingface.co:

https://huggingface.co/microsoft/VibeVoice-AcousticTokenizer

Url of VibeVoice-AcousticTokenizer

VibeVoice-AcousticTokenizer huggingface.co Url

Provider of VibeVoice-AcousticTokenizer huggingface.co

microsoft
ORGANIZATIONS

Other API from microsoft

huggingface.co

Total runs: 681.1K
Run Growth: 208.0K
Growth Rate: 30.53%
Updated:February 03 2022
huggingface.co

Total runs: 595.8K
Run Growth: -131.1K
Growth Rate: -22.00%
Updated:November 25 2025
huggingface.co

Total runs: 535.6K
Run Growth: 307.0K
Growth Rate: 57.32%
Updated:April 08 2024
huggingface.co

Total runs: 511.8K
Run Growth: -531.5K
Growth Rate: -103.84%
Updated:December 08 2025
huggingface.co

Total runs: 289.1K
Run Growth: -22.3K
Growth Rate: -7.70%
Updated:February 14 2024
huggingface.co

Total runs: 152.2K
Run Growth: -237.9K
Growth Rate: -156.31%
Updated:September 26 2022
huggingface.co

Total runs: 117.8K
Run Growth: -147.4K
Growth Rate: -125.12%
Updated:November 08 2023
huggingface.co

Total runs: 117.6K
Run Growth: -1.2K
Growth Rate: -1.03%
Updated:February 29 2024
huggingface.co

Total runs: 102.2K
Run Growth: -831
Growth Rate: -0.81%
Updated:February 03 2023
huggingface.co

Total runs: 100.0K
Run Growth: -425
Growth Rate: -0.42%
Updated:August 28 2025
huggingface.co

Total runs: 69.7K
Run Growth: -13.1K
Growth Rate: -18.74%
Updated:November 25 2025
huggingface.co

Total runs: 68.7K
Run Growth: -18.2K
Growth Rate: -26.47%
Updated:December 03 2025
huggingface.co

Total runs: 55.7K
Run Growth: -1.8K
Growth Rate: -3.23%
Updated:December 23 2021
huggingface.co

Total runs: 41.0K
Run Growth: -90.8K
Growth Rate: -221.30%
Updated:May 12 2026
huggingface.co

Total runs: 32.5K
Run Growth: 28.2K
Growth Rate: 86.85%
Updated:April 23 2026
huggingface.co

Total runs: 27.4K
Run Growth: -4.9K
Growth Rate: -18.05%
Updated:April 24 2023