Streaming, zero-shot, open-vocabulary keyword spotting for iOS / macOS / visionOS. Exported from icefall's KWS-finetuned Zipformer transducer (gigaspeech, 3.49M parameters) to CoreML with INT8 palettized weights, FP16 compute, and an iOS 17+ minimum deployment target.
Given an arbitrary list of English keywords at runtime (no retraining required), the model emits a match when it hears one. Runs ~26× real-time on Apple Silicon CPU + Neural Engine.
False positive rate: 0.27 / utterance on 60 random negative utterances. CoreML INT8 output agrees with the PyTorch FP32 reference on 99% of utterances (FP16 palettization drift is not systematic).
Note on "MISTER"
: SentencePiece tokenizes it as
[▁MI, S, TER]
(3 tokens). Stateless transducers rarely lock onto 3-token sequences in beam search. For production wake words, prefer single-token keywords.
Memory
Loaded all three
.mlmodelc
models: peak RSS delta ≈
63 MB
on macOS. After 218-utterance streaming workload: +127 MB total.
Usage
Python (coremltools)
import json
from pathlib import Path
import coremltools as ct
import numpy as np
model_dir = Path("./KWS-Zipformer-3M-CoreML-INT8")
config = json.loads((model_dir / "config.json").read_text())
encoder = ct.models.CompiledMLModel(str(model_dir / "encoder.mlmodelc"))
decoder = ct.models.CompiledMLModel(str(model_dir / "decoder.mlmodelc"))
joiner = ct.models.CompiledMLModel(str(model_dir / "joiner.mlmodelc"))
# Build zero state from config
state = {}
for name, shape inzip(config["encoder"]["layerStateNames"],
config["encoder"]["layerStateShapes"]):
state[name] = np.zeros(shape, dtype=np.float32)
state["cached_embed_left_pad"] = np.zeros(
config["encoder"]["cachedEmbedLeftPadShape"], dtype=np.float32)
state["processed_lens"] = np.zeros((1,), dtype=np.int32)
# Feed 45 mel frames (80-dim fbank) per chunk
x = np.zeros((1, 45, 80), dtype=np.float32) # replace with real fbank
out = encoder.predict({"x": x, **state})
encoder_out = out["encoder_out"] # (1, 8, 320) in joiner space# Stream: feed encoder_out[0, t] into decoder/joiner + beam search
Swift (speech-swift)
import SpeechSwift
let model =tryawaitKWSZipformerModel.fromPretrained(
"aufklarer/KWS-Zipformer-3M-CoreML-INT8"
)
let stream = model.streamingSession(keywords: ["HEY SONIQO", "STOP", "LIGHTS ON"])
try stream.feed(audio: pcm16k)
for match in stream.emissions {
print("matched: \(match.phrase) at \(match.timestamp)")
}
Reference Python decoder
The upstream Aho-Corasick ContextGraph + boost-cancellation beam search algorithm is available as a dependency-free pure-Python reference for porting to other runtimes.
Source
Exported from
icefall-kws-zipformer-gigaspeech-20240219
(Apache-2.0), specifically the
exp-finetune/pretrained.pt
KWS-finetuned checkpoint, published in
pkufool/keyword-spotting-models v0.11
.
KWS-Zipformer-3M-CoreML-INT8 huggingface.co is an AI model on huggingface.co that provides KWS-Zipformer-3M-CoreML-INT8's model effect (), which can be used instantly with this aufklarer KWS-Zipformer-3M-CoreML-INT8 model. huggingface.co supports a free trial of the KWS-Zipformer-3M-CoreML-INT8 model, and also provides paid use of the KWS-Zipformer-3M-CoreML-INT8. Support call KWS-Zipformer-3M-CoreML-INT8 model through api, including Node.js, Python, http.
KWS-Zipformer-3M-CoreML-INT8 huggingface.co is an online trial and call api platform, which integrates KWS-Zipformer-3M-CoreML-INT8's modeling effects, including api services, and provides a free online trial of KWS-Zipformer-3M-CoreML-INT8, you can try KWS-Zipformer-3M-CoreML-INT8 online for free by clicking the link below.
aufklarer KWS-Zipformer-3M-CoreML-INT8 online free url in huggingface.co:
KWS-Zipformer-3M-CoreML-INT8 is an open source model from GitHub that offers a free installation service, and any user can find KWS-Zipformer-3M-CoreML-INT8 on GitHub to install. At the same time, huggingface.co provides the effect of KWS-Zipformer-3M-CoreML-INT8 install, users can directly use KWS-Zipformer-3M-CoreML-INT8 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
KWS-Zipformer-3M-CoreML-INT8 install url in huggingface.co: