The "whisper-large-v3-tiny-caesar" is an acoustic model based on
"openai/whisper-large-v3"
suitable for Automatic Speech Recognition in code switching conditions between Spanish and Catalan.
Model Description
The "whisper-large-v3-tiny-caesar" is an acoustic model suitable for Automatic Speech Recognition in code switching conditions between Spanish and Catalan. It is the result of finetuning the model
"openai/whisper-large-v3"
with 2 hours of synthetic code switching data in Spanish/Catalan generated by the
Projecte AINA
from Barcelona, Spain.
CAESAR is an acronym with the following meaning:
(CA)talan (ES)panish (A)utomatic (R)ecognition
While "tiny" indicates that this model was finetuned with a very small amount of synthetic data (2 hours only).
Intended Uses and Limitations
This model can be used for Automatic Speech Recognition (ASR) in code switching conditions between Spanish and Catalan. The model is intended to transcribe audio files to plain text.
How to Get Started with the Model
To see an updated and functional version of this code, please see our our
Notebook
#This code works with GPU#Notice that: load_metric is no longer part of datasets.#you have to remove it and use evaluate's load instead.#(Note from November 2024)import torch
from transformers import WhisperForConditionalGeneration, WhisperProcessor
#Load the processor and model.
MODEL_NAME="projecte-aina/whisper-large-v3-tiny-caesar"
processor = WhisperProcessor.from_pretrained(MODEL_NAME)
model = WhisperForConditionalGeneration.from_pretrained(MODEL_NAME).to("cuda")
#Load the datasetfrom datasets import load_dataset, load_metric, Audio
ds=load_dataset("projecte-aina/3catparla_asr",split='test')
#Downsample to 16kHz
ds = ds.cast_column("audio", Audio(sampling_rate=16_000))
#Process the datasetdefmap_to_pred(batch):
audio = batch["audio"]
input_features = processor(audio["array"], sampling_rate=audio["sampling_rate"], return_tensors="pt").input_features
batch["reference"] = processor.tokenizer._normalize(batch['normalized_text'])
with torch.no_grad():
predicted_ids = model.generate(input_features.to("cuda"))[0]
transcription = processor.decode(predicted_ids)
batch["prediction"] = processor.tokenizer._normalize(transcription)
return batch
#Do the evaluation
result = ds.map(map_to_pred)
#Compute the overall WER now.from evaluate import load
wer = load("wer")
WER=100 * wer.compute(references=result["reference"], predictions=result["prediction"])
print(WER)
Training Details
Training data
The specific dataset used to create the model is a corpus called CAESAR-tiny which has not been released at the moment.
whisper-large-v3-tiny-caesar huggingface.co is an AI model on huggingface.co that provides whisper-large-v3-tiny-caesar's model effect (), which can be used instantly with this projecte-aina whisper-large-v3-tiny-caesar model. huggingface.co supports a free trial of the whisper-large-v3-tiny-caesar model, and also provides paid use of the whisper-large-v3-tiny-caesar. Support call whisper-large-v3-tiny-caesar model through api, including Node.js, Python, http.
whisper-large-v3-tiny-caesar huggingface.co is an online trial and call api platform, which integrates whisper-large-v3-tiny-caesar's modeling effects, including api services, and provides a free online trial of whisper-large-v3-tiny-caesar, you can try whisper-large-v3-tiny-caesar online for free by clicking the link below.
projecte-aina whisper-large-v3-tiny-caesar online free url in huggingface.co:
whisper-large-v3-tiny-caesar is an open source model from GitHub that offers a free installation service, and any user can find whisper-large-v3-tiny-caesar on GitHub to install. At the same time, huggingface.co provides the effect of whisper-large-v3-tiny-caesar install, users can directly use whisper-large-v3-tiny-caesar installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
whisper-large-v3-tiny-caesar install url in huggingface.co: