FluidInference / parakeet-ctc-110m-coreml

huggingface.co
Total runs: 120.3K
24-hour runs: 2.0K
7-day runs: 9.0K
30-day runs: 87.4K
Model's Last Updated: March 29 2026
automatic-speech-recognition

Introduction of parakeet-ctc-110m-coreml

Model Details of parakeet-ctc-110m-coreml

Parakeet-TDT-CTC-110M CoreML

NVIDIA's Parakeet-TDT-CTC-110M model converted to CoreML format for efficient inference on Apple Silicon.

Model Description

This is a hybrid ASR model with a shared Conformer encoder and two decoder heads:

  • CTC Head : Fast greedy decoding, ideal for keyword spotting
  • TDT Head : Token-Duration Transducer for high-quality transcription
Architecture
Component Description Size
Preprocessor Mel spectrogram extraction ~1 MB
Encoder Conformer encoder (shared) ~400 MB
CTCHead CTC output projection ~4 MB
Decoder TDT prediction network (LSTM) ~25 MB
JointDecision TDT joint network ~6 MB

Total size : ~436 MB

Performance

Benchmarked on Earnings22 dataset (772 audio files):

Metric Value
Keyword Recall 100% (1309/1309)
WER 17.97%
RTFx (M4 Pro) 358x real-time
Requirements
  • macOS 13+ (Ventura or later)
  • Apple Silicon (M1/M2/M3/M4)
  • Python 3.10+
Installation
# Using uv (recommended)
uv sync

# Or using pip
pip install -e .

# For audio file support (WAV, MP3, etc.)
pip install -e ".[audio]"
Usage
Python Inference
from scripts.inference import ParakeetCoreML

# Load model (from current directory with .mlpackage files)
model = ParakeetCoreML(".")

# Transcribe with TDT (higher quality)
text = model.transcribe("audio.wav", mode="tdt")
print(text)

# Or use CTC for faster keyword spotting
text = model.transcribe("audio.wav", mode="ctc")
print(text)
Command Line
# TDT decoding (default, higher quality)
uv run scripts/inference.py --audio audio.wav

# CTC decoding (faster, good for keyword spotting)
uv run scripts/inference.py --audio audio.wav --mode ctc
Model Conversion

To convert from the original NeMo model:

# Install conversion dependencies
uv sync --extra convert

# Run conversion
uv run scripts/convert_nemo_to_coreml.py --output-dir ./model

This will:

  1. Download the original model from NVIDIA ( nvidia/parakeet-tdt_ctc-110m )
  2. Convert each component to CoreML format
  3. Extract vocabulary and create metadata
File Structure
./
├── Preprocessor.mlpackage    # Audio → Mel spectrogram
├── Encoder.mlpackage         # Mel → Encoder features
├── CTCHead.mlpackage         # Encoder → CTC log probs
├── Decoder.mlpackage         # TDT prediction network
├── JointDecision.mlpackage   # TDT joint network
├── vocab.json                # Token vocabulary (1024 tokens)
├── metadata.json             # Model configuration
├── pyproject.toml            # Python dependencies
├── uv.lock                   # Locked dependencies
└── scripts/                  # Inference & conversion scripts
Decoding Modes
TDT Mode (Recommended for Transcription)
  • Uses Token-Duration Transducer decoding
  • Higher accuracy (17.97% WER)
  • Predicts both tokens and durations
  • Best for full transcription tasks
CTC Mode (Recommended for Keyword Spotting)
  • Greedy CTC decoding
  • Faster inference
  • 100% keyword recall on Earnings22
  • Best for detecting specific words/phrases
Custom Vocabulary / Keyword Spotting

For keyword spotting, CTC mode with custom vocabulary boosting achieves 100% recall:

# Load custom vocabulary with token IDs
with open("custom_vocab.json") as f:
    keywords = json.load(f)  # {"keyword": [token_ids], ...}

# Run CTC decoding
tokens = model.decode_ctc(encoder_output)

# Check for keyword matches
for keyword, expected_ids in keywords.items():
    if is_subsequence(expected_ids, tokens):
        print(f"Found keyword: {keyword}")
License

This model conversion is released under the Apache 2.0 License, same as the original NVIDIA model.

Citation

If you use this model, please cite the original NVIDIA work:

@misc{nvidia_parakeet_tdt_ctc,
  title={Parakeet-TDT-CTC-110M},
  author={NVIDIA},
  year={2024},
  publisher={Hugging Face},
  url={https://huggingface.co/nvidia/parakeet-tdt_ctc-110m}
}
Acknowledgments

Runs of FluidInference parakeet-ctc-110m-coreml on huggingface.co

120.3K
Total runs
2.0K
24-hour runs
5.9K
3-day runs
9.0K
7-day runs
87.4K
30-day runs

More Information About parakeet-ctc-110m-coreml huggingface.co Model

More parakeet-ctc-110m-coreml license Visit here:

https://choosealicense.com/licenses/cc-by-4.0

parakeet-ctc-110m-coreml huggingface.co

parakeet-ctc-110m-coreml huggingface.co is an AI model on huggingface.co that provides parakeet-ctc-110m-coreml's model effect (), which can be used instantly with this FluidInference parakeet-ctc-110m-coreml model. huggingface.co supports a free trial of the parakeet-ctc-110m-coreml model, and also provides paid use of the parakeet-ctc-110m-coreml. Support call parakeet-ctc-110m-coreml model through api, including Node.js, Python, http.

parakeet-ctc-110m-coreml huggingface.co Url

https://huggingface.co/FluidInference/parakeet-ctc-110m-coreml

FluidInference parakeet-ctc-110m-coreml online free

parakeet-ctc-110m-coreml huggingface.co is an online trial and call api platform, which integrates parakeet-ctc-110m-coreml's modeling effects, including api services, and provides a free online trial of parakeet-ctc-110m-coreml, you can try parakeet-ctc-110m-coreml online for free by clicking the link below.

FluidInference parakeet-ctc-110m-coreml online free url in huggingface.co:

https://huggingface.co/FluidInference/parakeet-ctc-110m-coreml

parakeet-ctc-110m-coreml install

parakeet-ctc-110m-coreml is an open source model from GitHub that offers a free installation service, and any user can find parakeet-ctc-110m-coreml on GitHub to install. At the same time, huggingface.co provides the effect of parakeet-ctc-110m-coreml install, users can directly use parakeet-ctc-110m-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

parakeet-ctc-110m-coreml install url in huggingface.co:

https://huggingface.co/FluidInference/parakeet-ctc-110m-coreml

Url of parakeet-ctc-110m-coreml

parakeet-ctc-110m-coreml huggingface.co Url

Provider of parakeet-ctc-110m-coreml huggingface.co

FluidInference
ORGANIZATIONS

Other API from FluidInference