aiola / whisper-medusa-v1

huggingface.co
Total runs: 70
24-hour runs: 8
7-day runs: 21
30-day runs: 61
Model's Last Updated: Agosto 04 2024

Introduction of whisper-medusa-v1

Model Details of whisper-medusa-v1

Whisper Medusa

Whisper is an advanced encoder-decoder model for speech transcription and translation, processing audio through encoding and decoding stages. Given its large size and slow inference speed, various optimization strategies like Faster-Whisper and Speculative Decoding have been proposed to enhance performance. Our Medusa model builds on Whisper by predicting multiple tokens per iteration, which significantly improves speed with small degradation in WER. We train and evaluate our model on the LibriSpeech dataset, demonstrating speed improvements.


Training Details

aiola/whisper-medusa-v1 was trained on the LibriSpeech dataset to perform audio translation. The Medusa heads were optimized for English, so for optimal performance and speed improvements, please use English audio only.


Usage

To use whisper-medusa-v1 install whisper-medusa repo following the README instructions.

Inference can be done using the following code:

import torch
import torchaudio

from whisper_medusa import WhisperMedusaModel
from transformers import WhisperProcessor

model_name = "aiola/whisper-medusa-v1"
model = WhisperMedusaModel.from_pretrained(model_name)
processor = WhisperProcessor.from_pretrained(model_name)

path_to_audio = "path/to/audio.wav"
SAMPLING_RATE = 16000
language = "en"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

input_speech, sr = torchaudio.load(path_to_audio)
if sr != SAMPLING_RATE:
    input_speech = torchaudio.transforms.Resample(sr, SAMPLING_RATE)(input_speech)

input_features = processor(input_speech.squeeze(), return_tensors="pt", sampling_rate=SAMPLING_RATE).input_features
input_features = input_features.to(device)

model = model.to(device)
model_output = model.generate(
    input_features,
    language=language,
)
predict_ids = model_output[0]
pred = processor.decode(predict_ids, skip_special_tokens=True)
print(pred)

Runs of aiola whisper-medusa-v1 on huggingface.co

70
Total runs
8
24-hour runs
13
3-day runs
21
7-day runs
61
30-day runs

More Information About whisper-medusa-v1 huggingface.co Model

More whisper-medusa-v1 license Visit here:

https://choosealicense.com/licenses/mit

whisper-medusa-v1 huggingface.co

whisper-medusa-v1 huggingface.co is an AI model on huggingface.co that provides whisper-medusa-v1's model effect (), which can be used instantly with this aiola whisper-medusa-v1 model. huggingface.co supports a free trial of the whisper-medusa-v1 model, and also provides paid use of the whisper-medusa-v1. Support call whisper-medusa-v1 model through api, including Node.js, Python, http.

whisper-medusa-v1 huggingface.co Url

https://huggingface.co/aiola/whisper-medusa-v1

aiola whisper-medusa-v1 online free

whisper-medusa-v1 huggingface.co is an online trial and call api platform, which integrates whisper-medusa-v1's modeling effects, including api services, and provides a free online trial of whisper-medusa-v1, you can try whisper-medusa-v1 online for free by clicking the link below.

aiola whisper-medusa-v1 online free url in huggingface.co:

https://huggingface.co/aiola/whisper-medusa-v1

whisper-medusa-v1 install

whisper-medusa-v1 is an open source model from GitHub that offers a free installation service, and any user can find whisper-medusa-v1 on GitHub to install. At the same time, huggingface.co provides the effect of whisper-medusa-v1 install, users can directly use whisper-medusa-v1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

whisper-medusa-v1 install url in huggingface.co:

https://huggingface.co/aiola/whisper-medusa-v1

Url of whisper-medusa-v1

whisper-medusa-v1 huggingface.co Url

Provider of whisper-medusa-v1 huggingface.co

aiola
ORGANIZATIONS

Other API from aiola

huggingface.co

Total runs: 128
Run Growth: 8
Growth Rate: 6.25%
Updated:Noviembre 22 2024
huggingface.co

Total runs: 87
Run Growth: 57
Growth Rate: 65.52%
Updated:Febrero 26 2026