The model can be used directly (without a language model) as follows:
import torch
import torchaudio
from datasets import load_dataset
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import re
test_dataset = load_dataset("common_voice", "br", split="test[:2%]")
processor = Wav2Vec2Processor.from_pretrained("cahya/wav2vec2-large-xlsr-breton")
model = Wav2Vec2ForCTC.from_pretrained("cahya/wav2vec2-large-xlsr-breton")
chars_to_ignore_regex = '[\\,\,\?\.\!\;\:\"\“\%\”\�\(\)\/\«\»\½\…]'# Preprocessing the datasets.# We need to read the aduio files as arraysdefspeech_file_to_array_fn(batch):
batch["sentence"] = re.sub(chars_to_ignore_regex, '', batch["sentence"]).lower() + " "
batch["sentence"] = batch["sentence"].replace("ʼ", "'")
batch["sentence"] = batch["sentence"].replace("’", "'")
batch["sentence"] = batch["sentence"].replace('‘', "'")
speech_array, sampling_rate = torchaudio.load(batch["path"])
resampler = torchaudio.transforms.Resample(sampling_rate, 16_000)
batch["speech"] = resampler(speech_array).squeeze().numpy()
return batch
test_dataset = test_dataset.map(speech_file_to_array_fn)
inputs = processor(test_dataset[:2]["speech"], sampling_rate=16_000, return_tensors="pt", padding=True)
with torch.no_grad():
logits = model(inputs.input_values, attention_mask=inputs.attention_mask).logits
predicted_ids = torch.argmax(logits, dim=-1)
print("Prediction:", processor.batch_decode(predicted_ids))
print("Reference:", test_dataset[:2]["sentence"])
The above code leads to the following prediction for the first two samples:
Prediction: ["ne' ler ket don a-benn us netra pa vez zer nic'hed evel-si", 'an eil hag egile']
Reference: ['"n\'haller ket dont a-benn eus netra pa vezer nec\'het evel-se." ', 'an eil hag egile. ']
Evaluation
The model can be evaluated as follows on the Breton test data of Common Voice.
import torch
import torchaudio
from datasets import load_dataset, load_metric
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import re
test_dataset = load_dataset("common_voice", "br", split="test")
wer = load_metric("wer")
processor = Wav2Vec2Processor.from_pretrained("cahya/wav2vec2-large-xlsr-breton")
model = Wav2Vec2ForCTC.from_pretrained("cahya/wav2vec2-large-xlsr-breton")
model.to("cuda")
chars_to_ignore_regex = '[\\,\,\?\.\!\;\:\"\“\%\”\�\(\)\/\«\»\½\…]'# Preprocessing the datasets.# We need to read the aduio files as arraysdefspeech_file_to_array_fn(batch):
batch["sentence"] = re.sub(chars_to_ignore_regex, '', batch["sentence"]).lower() + " "
batch["sentence"] = batch["sentence"].replace("ʼ", "'")
batch["sentence"] = batch["sentence"].replace("’", "'")
batch["sentence"] = batch["sentence"].replace('‘', "'")
speech_array, sampling_rate = torchaudio.load(batch["path"])
resampler = torchaudio.transforms.Resample(sampling_rate, 16_000)
batch["speech"] = resampler(speech_array).squeeze().numpy()
return batch
test_dataset = test_dataset.map(speech_file_to_array_fn)
# Preprocessing the datasets.# We need to read the aduio files as arraysdefevaluate(batch):
inputs = processor(batch["speech"], sampling_rate=16_000, return_tensors="pt", padding=True)
with torch.no_grad():
logits = model(inputs.input_values.to("cuda"), attention_mask=inputs.attention_mask.to("cuda")).logits
pred_ids = torch.argmax(logits, dim=-1)
batch["pred_strings"] = processor.batch_decode(pred_ids)
return batch
result = test_dataset.map(evaluate, batched=True, batch_size=8)
print("WER: {:2f}".format(100 * wer.compute(predictions=result["pred_strings"], references=result["sentence"])))
Test Result
: 41.71 %
Training
The Common Voice
train
,
validation
, and ... datasets were used for training as well as ... and ... # TODO
The script used for training can be found
here
(will be available soon)
Runs of cahya wav2vec2-large-xlsr-breton on huggingface.co
26
Total runs
0
24-hour runs
3
3-day runs
6
7-day runs
5
30-day runs
More Information About wav2vec2-large-xlsr-breton huggingface.co Model
More wav2vec2-large-xlsr-breton license Visit here:
wav2vec2-large-xlsr-breton huggingface.co is an AI model on huggingface.co that provides wav2vec2-large-xlsr-breton's model effect (), which can be used instantly with this cahya wav2vec2-large-xlsr-breton model. huggingface.co supports a free trial of the wav2vec2-large-xlsr-breton model, and also provides paid use of the wav2vec2-large-xlsr-breton. Support call wav2vec2-large-xlsr-breton model through api, including Node.js, Python, http.
wav2vec2-large-xlsr-breton huggingface.co is an online trial and call api platform, which integrates wav2vec2-large-xlsr-breton's modeling effects, including api services, and provides a free online trial of wav2vec2-large-xlsr-breton, you can try wav2vec2-large-xlsr-breton online for free by clicking the link below.
cahya wav2vec2-large-xlsr-breton online free url in huggingface.co:
wav2vec2-large-xlsr-breton is an open source model from GitHub that offers a free installation service, and any user can find wav2vec2-large-xlsr-breton on GitHub to install. At the same time, huggingface.co provides the effect of wav2vec2-large-xlsr-breton install, users can directly use wav2vec2-large-xlsr-breton installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
wav2vec2-large-xlsr-breton install url in huggingface.co: