We introduce WhisperNER, a novel model that allows joint speech transcription and entity recognition.
WhisperNER supports open-type NER, enabling recognition of diverse and evolving entities at inference.
Training Details
aiola/whisper-ner-v1
was trained on the NuNER dataset to perform joint audio transcription and NER tagging.
The model was trained and evaluated only on English data. Check out the
paper
for full details.
Usage
Inference can be done using the following code (for inference code and more details check out the
whisper-ner repo
).:
import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration
model_path = "aiola/whisper-ner-v1"
audio_file_path = "path/to/audio/file"
prompt = "person, company, location"# comma separated entity tags# load model and processor from pre-trained
processor = WhisperProcessor.from_pretrained(model_path)
model = WhisperForConditionalGeneration.from_pretrained(model_path)
device = torch.device("cuda"if torch.cuda.is_available() else"cpu")
model = model.to(device)
# load audio file: user is responsible for loading the audio files themselves
target_sample_rate = 16000
signal, sampling_rate = torchaudio.load(audio_file_path)
resampler = torchaudio.transforms.Resample(sampling_rate, target_sample_rate)
signal = resampler(signal)
# convert to mono or remove first dim if neededif signal.ndim == 2:
signal = torch.mean(signal, dim=0)
# pre-process to get the input features
input_features = processor(
signal, sampling_rate=target_sample_rate, return_tensors="pt"
).input_features
input_features = input_features.to(device)
prompt_ids = processor.get_prompt_ids(prompt.lower(), return_tensors="pt")
prompt_ids = prompt_ids.to(device)
# generate token ids by running model forward sequentiallywith torch.no_grad():
predicted_ids = model.generate(
input_features,
prompt_ids=prompt_ids,
generation_config=model.generation_config,
language="en",
)
# post-process token ids to text, remove prompt
transcription = processor.batch_decode(
predicted_ids[:, prompt_ids.shape[0]:], skip_special_tokens=True
)[0]
print(transcription)
Runs of aiola whisper-ner-v1 on huggingface.co
89
Total runs
0
24-hour runs
7
3-day runs
-22
7-day runs
-24
30-day runs
More Information About whisper-ner-v1 huggingface.co Model
whisper-ner-v1 huggingface.co is an AI model on huggingface.co that provides whisper-ner-v1's model effect (), which can be used instantly with this aiola whisper-ner-v1 model. huggingface.co supports a free trial of the whisper-ner-v1 model, and also provides paid use of the whisper-ner-v1. Support call whisper-ner-v1 model through api, including Node.js, Python, http.
whisper-ner-v1 huggingface.co is an online trial and call api platform, which integrates whisper-ner-v1's modeling effects, including api services, and provides a free online trial of whisper-ner-v1, you can try whisper-ner-v1 online for free by clicking the link below.
aiola whisper-ner-v1 online free url in huggingface.co:
whisper-ner-v1 is an open source model from GitHub that offers a free installation service, and any user can find whisper-ner-v1 on GitHub to install. At the same time, huggingface.co provides the effect of whisper-ner-v1 install, users can directly use whisper-ner-v1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.