A transformer encoder-decoder model for automatic audio captioning. As opposed to speech-to-text, captioning describes the content of audio clips, such as prominent sounds or environmental noises. This task has numerous practical applications, e.g., for providing access to audio information for people with hearing impairments or improving the searchability of audio content.
Example output:
clotho > caption: Rain is pouring down and thunder is rumbling in the background.
The style prefix influences the style of the caption. Model knows 3 styles:
audioset > keywords:
,
audiocaps > caption:
, and
clotho > caption:
. It was finetuned on Clotho and that is the indended "default" style.
WhisperTokenizer must be initialized with
language="en"
and
task="transcribe"
.
Our model class
WhisperForAudioCaptioning
can be found in our git repository or here on the HuggingFace Hub in the model repository. The class overrides default Whisper
generate
method to support forcing decoder prefix.
Training details
The model was initialized by original speech-to-text
openai/whisper-small
weights. Then, it was pretrained on a mix of (1) subset of AudioSet with synthetic labels, (2) AudioCaps captioning dataset and (3) Clotho v2.1 captioning dataset. Finally, it was finetuned on Clotho v2.1 to focus the model on the specific style of the captions. For each traning input, the model was informed about the source of the data, so it can mimic the caption style in all 3 styles.
During pretraining, the ratio of samples in each batch was approximately 12:3:1 (AudioSet:AudioCaps:Clotho). The pretraining took 19800 steps with batch size 32 and learning rate 2e-5. Finetuning was done on Clotho only, and the model was trained for 1500 steps with batch size 32 and learning rate 4e-6. All layers except
fc1
layers were frozen during finetuning.
For more information about the training regime, see the
technical report
.
Evaluation details
Metrics reported in the metadata were computed on Clotho v2.1 test split with captions generated using a beam search with 5 beams.
whisper-tiny
whisper-small
whisper-large-v2
SacreBLEU
13.77
15.76
16.50
METEOR
0.3452
0.3781
0.3782
CIDEr
0.3404
0.4142
0.4331
SPICE
0.1077
0.1234
0.1257
SPIDEr
0.2240
0.2687
0.2794
Limitations
The captions generated by the model can be misleading or not truthful, even if they appear convincing. The hallucination occurs especially in domains that were not present in the finetuning data.
While the original speech-to-text checkpoints by OpenAI were trained on multilingual data, our training contains only English captions, and therefore is not expected for the model to support other languages.
Licence
The model weights are published under non-commercial license CC BY-NC 4.0 as the model was finetuned on a dataset for non-commercial use.
Contact
If you'd like to chat about this, please get in touch with is via email at kadlcik
<at>
mail.muni.cz or ahajek
<at>
mail.muni.cz.
Runs of MU-NLPC whisper-small-audio-captioning on huggingface.co
298
Total runs
1
24-hour runs
14
3-day runs
43
7-day runs
188
30-day runs
More Information About whisper-small-audio-captioning huggingface.co Model
More whisper-small-audio-captioning license Visit here:
whisper-small-audio-captioning huggingface.co is an AI model on huggingface.co that provides whisper-small-audio-captioning's model effect (), which can be used instantly with this MU-NLPC whisper-small-audio-captioning model. huggingface.co supports a free trial of the whisper-small-audio-captioning model, and also provides paid use of the whisper-small-audio-captioning. Support call whisper-small-audio-captioning model through api, including Node.js, Python, http.
whisper-small-audio-captioning huggingface.co is an online trial and call api platform, which integrates whisper-small-audio-captioning's modeling effects, including api services, and provides a free online trial of whisper-small-audio-captioning, you can try whisper-small-audio-captioning online for free by clicking the link below.
MU-NLPC whisper-small-audio-captioning online free url in huggingface.co:
whisper-small-audio-captioning is an open source model from GitHub that offers a free installation service, and any user can find whisper-small-audio-captioning on GitHub to install. At the same time, huggingface.co provides the effect of whisper-small-audio-captioning install, users can directly use whisper-small-audio-captioning installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
whisper-small-audio-captioning install url in huggingface.co: