Benchmark on FluidInference/fleurs-full (650 Japanese samples)
:
CER
: 10.29% (within expected 10-13% range)
RTFx
: 136.85x (far exceeds real-time)
Avg Latency
: 91.34ms per sample on M-series chips
Expected CER by Dataset
(from NeMo paper):
Dataset
CER
JSUT basic5000
6.5%
Mozilla Common Voice 8.0 test
7.2%
Mozilla Common Voice 16.1 dev
10.2%
Mozilla Common Voice 16.1 test
13.3%
TEDxJP-10k
9.1%
Critical Implementation Note: Raw Logits Output
IMPORTANT
: The CTC decoder outputs
raw logits
(not log-probabilities). You
must
apply
log_softmax
before CTC decoding.
Why?
During CoreML conversion, we discovered that
log_softmax
failed to convert correctly, producing extreme values (-45440 instead of -67). The solution was to output raw logits and apply
log_softmax
in post-processing.
Usage Example
import coremltools as ct
import numpy as np
import torch
# Load the three CoreML models
preprocessor = ct.models.MLModel('Preprocessor.mlpackage')
encoder = ct.models.MLModel('Encoder.mlpackage')
ctc_decoder = ct.models.MLModel('CtcDecoder.mlpackage')
# Prepare audio (16kHz, mono, max 15 seconds)
audio = np.array(audio_samples, dtype=np.float32).reshape(1, -1)
audio_length = np.array([audio.shape[1]], dtype=np.int32)
# Pad or truncate to 240,000 samples (15 seconds)if audio.shape[1] < 240000:
audio = np.pad(audio, ((0, 0), (0, 240000 - audio.shape[1])))
else:
audio = audio[:, :240000]
# Step 1: Preprocessor (audio → mel)
prep_out = preprocessor.predict({
'audio_signal': audio,
'length': audio_length
})
# Step 2: Encoder (mel → features)
enc_out = encoder.predict({
'mel_features': prep_out['mel_features'],
'mel_length': prep_out['mel_length']
})
# Step 3: CTC Decoder (features → raw logits)
ctc_out = ctc_decoder.predict({
'encoder_output': enc_out['encoder_output']
})
raw_logits = ctc_out['ctc_logits'] # [1, 188, 3073]# Apply log_softmax (CRITICAL!)
logits_tensor = torch.from_numpy(raw_logits)
log_probs = torch.nn.functional.log_softmax(logits_tensor, dim=-1)
# Now use log_probs for CTC decoding# Greedy decoding example:
labels = torch.argmax(log_probs, dim=-1)[0].numpy() # [188]# Collapse repeats and remove blanks
blank_id = 3072
decoded = []
prev = Nonefor label in labels:
if label != blank_id and label != prev:
decoded.append(label)
prev = label
# Convert to text using vocabularyimport json
withopen('vocab.json', 'r') as f:
vocab = json.load(f)
tokens = [vocab[i] for i in decoded if i < len(vocab)]
text = ''.join(tokens).replace('▁', ' ').strip()
print(text)
parakeet-ctc-0.6b-ja-coreml huggingface.co is an AI model on huggingface.co that provides parakeet-ctc-0.6b-ja-coreml's model effect (), which can be used instantly with this FluidInference parakeet-ctc-0.6b-ja-coreml model. huggingface.co supports a free trial of the parakeet-ctc-0.6b-ja-coreml model, and also provides paid use of the parakeet-ctc-0.6b-ja-coreml. Support call parakeet-ctc-0.6b-ja-coreml model through api, including Node.js, Python, http.
parakeet-ctc-0.6b-ja-coreml huggingface.co is an online trial and call api platform, which integrates parakeet-ctc-0.6b-ja-coreml's modeling effects, including api services, and provides a free online trial of parakeet-ctc-0.6b-ja-coreml, you can try parakeet-ctc-0.6b-ja-coreml online for free by clicking the link below.
FluidInference parakeet-ctc-0.6b-ja-coreml online free url in huggingface.co:
parakeet-ctc-0.6b-ja-coreml is an open source model from GitHub that offers a free installation service, and any user can find parakeet-ctc-0.6b-ja-coreml on GitHub to install. At the same time, huggingface.co provides the effect of parakeet-ctc-0.6b-ja-coreml install, users can directly use parakeet-ctc-0.6b-ja-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
parakeet-ctc-0.6b-ja-coreml install url in huggingface.co: