Respair / Whisper_JP

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: August 12 2024

Introduction of Whisper_JP

Model Details of Whisper_JP

Hibiki ASR Phonemizer

This model is a Phoneme Level Speech Recognition network, originally a fine-tuned version of openai/whisper-large-v3 on a mixture of Different Japanese datasets.

it can detect, transcribe and do the following:

  • non-speech sounds such as gasp, erotic moans, laughter, etc.
  • adding punctuations more faithfully.

a Grapheme level language modelling head (i.e outputting normal Japanese) will probably be trained as well. Though going directly from audio to Phonemes will result in a more accurate representation for Japnese.

evaluation set:

  • Loss: 0.2186
  • Wer: 21.6707
Inference and Post-proc

# this function was borrowed and modified from Aaron Yinghao Li, the Author of StyleTTS paper.

from datasets import Dataset, Audio
from transformers import WhisperProcessor, WhisperForConditionalGeneration
import jaconv

kana_mapper = dict([
    ("ゔぁ","ba"),
          .
          .
          .
          etc. # Take a look at the Notebook for the whole code
    ("ぉ"," o"),
    ("ゎ"," ɯa"),
    ("ぉ"," o"),

    ("を","o")
])


def post_fix(text):
    orig = text

    for k, v in kana_mapper.items():
        text = text.replace(k, v)

    return text


processor = WhisperProcessor.from_pretrained("openai/whisper-large-v3")
model = WhisperForConditionalGeneration.from_pretrained("Respair/Hibiki_ASR_Phonemizer").to("cuda:0")

forced_decoder_ids = processor.get_decoder_prompt_ids(task="transcribe", language='japanese')


import re

sample = Dataset.from_dict({"audio": ["/content/kl_chunk1987.wav"]}).cast_column("audio", Audio(16000))
sample = sample[0]['audio']

# Ensure the input features are on the same device as the model
input_features = processor(sample["array"], sampling_rate=sample["sampling_rate"], return_tensors="pt").input_features.to("cuda:0")

# generate token ids
predicted_ids = model.generate(input_features,forced_decoder_ids=forced_decoder_ids, repetition_penalty=1.2)
# decode token ids to text
transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)


# You can add your final adjustments here, it's better to write a dict though, but I'm just giving you a quick demonstration here.

if ' neɽitai ' in transcription[0]:
    transcription[0] = transcription[0].replace(' neɽitai ', "naɽitai")

if 'harɯdʑisama' in transcription[0]:
    transcription[0] = transcription[0].replace('harɯdʑisama', "arɯdʑisama")


if "ki ni ɕinai" in transcription[0]:
    transcription[0] = re.sub(r'(?<!\s)ki ni ɕinai', r' ki ni ɕinai', transcription[0])

if 'ʔt' in transcription[0]:
    transcription[0] = re.sub(r'(?<!\s)ʔt', r'ʔt', transcription[0])

if 'de aɽoɯ' in transcription[0]:
    transcription[0] = re.sub(r'(?<!\s)de aɽoɯ', r' de aɽoɯ', transcription[0])

post_fix(jaconv.kata2hira(transcription[0].lstrip())) # Ensuring the model won't hallucinate and return kana

the Full code -> Notebook

Intended uses & limitations

No restrictions is imposed by me, but proceed at your own risk, The User (You) are entirely responisble for their actions.

Training and evaluation data
  • Japanese Common Voice 17
  • ehehe Corpus
  • Custom Game and Anime dataset (around 8 hours)
Training procedure
Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-05
  • train_batch_size: 24
  • eval_batch_size: 8
  • seed: 42
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 500
  • training_steps: 6000
Training results
Training Loss Epoch Step Validation Loss Wer
0.2101 0.8058 1000 0.2090 30.1840
0.1369 1.6116 2000 0.1837 27.6756
0.0838 2.4174 3000 0.1829 26.4036
0.0454 3.2232 4000 0.1922 20.9549
0.0434 4.0290 5000 0.2072 20.8898
0.021 4.8348 6000 0.2186 21.6707
Compute and Duration
  • 1x A100(40G)
  • 64gb RAM
  • BF16
  • 14hrs
Framework versions
  • Transformers 4.41.1
  • Pytorch 2.4.0+cu121
  • Datasets 2.19.1
  • Tokenizers 0.19.1

Runs of Respair Whisper_JP on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About Whisper_JP huggingface.co Model

More Whisper_JP license Visit here:

https://choosealicense.com/licenses/apache-2.0

Whisper_JP huggingface.co

Whisper_JP huggingface.co is an AI model on huggingface.co that provides Whisper_JP's model effect (), which can be used instantly with this Respair Whisper_JP model. huggingface.co supports a free trial of the Whisper_JP model, and also provides paid use of the Whisper_JP. Support call Whisper_JP model through api, including Node.js, Python, http.

Whisper_JP huggingface.co Url

https://huggingface.co/Respair/Whisper_JP

Respair Whisper_JP online free

Whisper_JP huggingface.co is an online trial and call api platform, which integrates Whisper_JP's modeling effects, including api services, and provides a free online trial of Whisper_JP, you can try Whisper_JP online for free by clicking the link below.

Respair Whisper_JP online free url in huggingface.co:

https://huggingface.co/Respair/Whisper_JP

Whisper_JP install

Whisper_JP is an open source model from GitHub that offers a free installation service, and any user can find Whisper_JP on GitHub to install. At the same time, huggingface.co provides the effect of Whisper_JP install, users can directly use Whisper_JP installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Whisper_JP install url in huggingface.co:

https://huggingface.co/Respair/Whisper_JP

Url of Whisper_JP

Whisper_JP huggingface.co Url

Provider of Whisper_JP huggingface.co

Respair
ORGANIZATIONS

Other API from Respair

huggingface.co

Total runs: 50
Run Growth: 40
Growth Rate: 80.00%
Updated:October 07 2025
huggingface.co

Total runs: 4
Run Growth: 4
Growth Rate: 100.00%
Updated:September 13 2025
huggingface.co

Total runs: 3
Run Growth: 0
Growth Rate: 0.00%
Updated:November 06 2024
huggingface.co

Total runs: 2
Run Growth: 2
Growth Rate: 100.00%
Updated:September 20 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 29 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:June 06 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:June 06 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:October 22 2024
huggingface.co

Total runs: 0
Run Growth: -6
Growth Rate: 0.00%
Updated:April 16 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 30 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:November 17 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 25 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 01 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:February 17 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 31 2025