Enhanced Silero VAD v3 model with trained Squeeze-Excitation (SE) modules for improved voice activity detection on Apple platforms.
Model Description
This is an enhanced version of the Silero VAD v3 model, converted to CoreML format with additional SE modules trained on the MUSAN dataset. The model provides state-of-the-art voice activity detection performance optimized for Apple Silicon.
Key Features
Single unified model
replacing the previous 3-model pipeline
Trained SE modules
for better noise/music suppression
92% accuracy
on MUSAN test set (F1-score: 91.3%)
stateless
we will add a stateful version in future uploads for 10%+ improvements
Architecture
The model uses the Silero VAD v3 architecture with the following enhancements:
LSTM-based encoder with 4 blocks
Squeeze-Excitation modules in each encoder block (trained on MUSAN)
LayerNorm after LSTM with 0.15 scaling factor
Input: 512 audio samples at 16kHz (32ms chunks)
Output: Voice probability [0, 1]
Performance
Based on 100 files, 50 noise, 50 speech
Metric
Value
Accuracy
92%
Precision
100%
Recall
84%
F1-Score
91.3%
RTFx
117-140x
Model Size
~1.5 MB
Files
silero_vad.mlmodelc
- Compiled CoreML model (for production)
Usage
Swift Integration
import CoreML
// Load the modellet modelURL =Bundle.main.url(forResource: "silero_vad", withExtension: "mlmodelc")!let model =tryMLModel(contentsOf: modelURL)
// Prepare input (512 samples at 16kHz)let audioArray =tryMLMultiArray(shape: [1, 512], dataType: .float32)
// ... fill audioArray with normalized audio samples [-1, 1]// Run inferencelet input =tryMLDictionaryFeatureProvider(dictionary: ["audio_chunk": audioArray])
let output = model.prediction(from: input)
let probability = output.featureValue(for: "vad_probability")!.multiArrayValue![0].floatValue
// Apply thresholdlet isVoiceActive = probability >=0.5
Important Notes
Audio Normalization
: Input audio must be normalized to [-1, 1] range
Chunk Size
: Fixed at 512 samples (32ms at 16kHz)
Threshold
: Recommended threshold is 0.5 (adjustable based on use case)
Training Details
The SE modules were trained on the MUSAN dataset with:
86.47% validation accuracy on MUSAN
Batch size: 32
Learning rate: 1e-3
Training samples: ~45,000 chunks
Optimizer: Adam
Requirements
macOS 13.0+ / iOS 16.0+
CoreML framework
Apple Silicon recommended for optimal performance
License
Model weights derived from Silero VAD v3 (MIT License).
SE module enhancements and CoreML conversion by FluidInference.
Runs of FluidInference silero-vad-v2-coreml on huggingface.co
10
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
1
30-day runs
More Information About silero-vad-v2-coreml huggingface.co Model
silero-vad-v2-coreml huggingface.co
silero-vad-v2-coreml huggingface.co is an AI model on huggingface.co that provides silero-vad-v2-coreml's model effect (), which can be used instantly with this FluidInference silero-vad-v2-coreml model. huggingface.co supports a free trial of the silero-vad-v2-coreml model, and also provides paid use of the silero-vad-v2-coreml. Support call silero-vad-v2-coreml model through api, including Node.js, Python, http.
silero-vad-v2-coreml huggingface.co is an online trial and call api platform, which integrates silero-vad-v2-coreml's modeling effects, including api services, and provides a free online trial of silero-vad-v2-coreml, you can try silero-vad-v2-coreml online for free by clicking the link below.
FluidInference silero-vad-v2-coreml online free url in huggingface.co:
silero-vad-v2-coreml is an open source model from GitHub that offers a free installation service, and any user can find silero-vad-v2-coreml on GitHub to install. At the same time, huggingface.co provides the effect of silero-vad-v2-coreml install, users can directly use silero-vad-v2-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
silero-vad-v2-coreml install url in huggingface.co: