This model is finetuned on top of feature extractor
VoxRex-model
from the National Library of Sweden. The finetuned model achieves the following results on the test set with a 5-gram KenLM. The numbers in parentheses are the results without the language model:
WER: 0.0703
(0.0979)
CER: 0.0269
(0.0311)
Model description
This is one of several Wav2Vec-models our team created during the 🤗 hosted
Robust Speech Event
. This is the complete list of our models and their final scores:
We have released all the code developed during the event so that the Norwegian NLP community can build upon it when developing even better Norwegian ASR models. The finetuning of these models is not very computationally demanding. After following the instructions here, you should be able to train your own automatic speech recognition system in less than a day with an average GPU.
Team
The following people contributed to building this model: Rolv-Arild Braaten, Per Egil Kummervold, Andre Kåsen, Javier de la Rosa, Per Erik Solberg, and Freddy Wetjen.
Training procedure
To reproduce these results, we strongly recommend that you follow the
instructions from 🤗
to train a simple Swedish model.
When you have verified that you are able to do this, create a fresh new repo. You can then start by copying the files
run.sh
and
run_speech_recognition_ctc.py
from our repo. Running these will create all the other necessary files, and should let you reproduce our results. With some tweaks to the hyperparameters, you might even be able to build an even better ASR. Good luck!
Language Model
As the scores indicate, adding even a simple 5-gram language will improve the results. 🤗 has provided another
very nice blog
explaining how to add a 5-gram language model to improve the ASR model. You can build this from your own corpus, for instance by extracting some suitable text from the
Norwegian Colossal Corpus
. You can also skip some of the steps in the guide, and copy the
5-gram model from this repo
.
Using these settings, the training might take 3-4 days on an average GPU. You can, however, get a decent model and faster results by tweaking these parameters.
Parameter
Comment
per_device_train_batch_size
Adjust this to the maximum of available memory. 16 or 24 might be good settings depending on your system
gradient_accumulation_steps
Can be adjusted even further up to increase batch size and speed up training without running into memory issues
learning_rate
Can be increased, maybe as high as 1e-4. Speeds up training but might add instability
epochs
Can be decreased significantly. This is a huge dataset and you might get a decent result already after a couple of epochs
Citation
@inproceedings{de-la-rosa-etal-2023-boosting,
title = "Boosting {N}orwegian Automatic Speech Recognition",
author = "De La Rosa, Javier and
Braaten, Rolv-Arild and
Kummervold, Per and
Wetjen, Freddy",
booktitle = "Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)",
month = may,
year = "2023",
address = "T{\'o}rshavn, Faroe Islands",
publisher = "University of Tartu Library",
url = "https://aclanthology.org/2023.nodalida-1.55",
pages = "555--564",
abstract = "In this paper, we present several baselines for automatic speech recognition (ASR) models for the two official written languages in Norway: Bokm{\aa}l and Nynorsk. We compare the performance of models of varying sizes and pre-training approaches on multiple Norwegian speech datasets. Additionally, we measure the performance of these models against previous state-of-the-art ASR models, as well as on out-of-domain datasets. We improve the state of the art on the Norwegian Parliamentary Speech Corpus (NPSC) from a word error rate (WER) of 17.10{\%} to 7.60{\%}, with models achieving 5.81{\%} for Bokm{\aa}l and 11.54{\%} for Nynorsk. We also discuss the challenges and potential solutions for further improving ASR models for Norwegian.",
}
nb-wav2vec2-300m-bokmaal huggingface.co is an AI model on huggingface.co that provides nb-wav2vec2-300m-bokmaal's model effect (), which can be used instantly with this NbAiLab nb-wav2vec2-300m-bokmaal model. huggingface.co supports a free trial of the nb-wav2vec2-300m-bokmaal model, and also provides paid use of the nb-wav2vec2-300m-bokmaal. Support call nb-wav2vec2-300m-bokmaal model through api, including Node.js, Python, http.
nb-wav2vec2-300m-bokmaal huggingface.co is an online trial and call api platform, which integrates nb-wav2vec2-300m-bokmaal's modeling effects, including api services, and provides a free online trial of nb-wav2vec2-300m-bokmaal, you can try nb-wav2vec2-300m-bokmaal online for free by clicking the link below.
NbAiLab nb-wav2vec2-300m-bokmaal online free url in huggingface.co:
nb-wav2vec2-300m-bokmaal is an open source model from GitHub that offers a free installation service, and any user can find nb-wav2vec2-300m-bokmaal on GitHub to install. At the same time, huggingface.co provides the effect of nb-wav2vec2-300m-bokmaal install, users can directly use nb-wav2vec2-300m-bokmaal installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
nb-wav2vec2-300m-bokmaal install url in huggingface.co: