PASE is a state-of-the-art generative speech enhancement model trained to remove noise and reverberation while preserving linguistic content and speaker identity. It operates on 16 kHz mono audio.
Model Details
Model Description
PASE contains two main components:
Denoising WavLM (DeWavLM)
Fine‑tuned from WavLM‑Large using denoising representation distillation (DRD).
Performs robust noise supression while effectively mitigating linguistic hallucinations by leveraging the phonological prior from self-supervised WavLM.
Dual‑Stream Vocoder
Reconstructs audio using DeWavLM's dual-stream representations:
Phonetic representation
: high-level linguistic structure
Acoustic representation
: speaker identity and prosody
These source datasets were used to prepare training mixtures and train the released model. The model card and repository do not redistribute the underlying dataset contents; please refer to the original dataset pages and licenses below.
Dataset Attribution
DNS5 Challenge clean speech (LibriVox subset): clean-speech material prepared from
LibriVox
through the
DNS Challenge
. The LibriVox recordings used for this portion are
public domain
and were used as clean-speech training data for the released checkpoint.
LibriTTS:
LibriTTS
by Heiga Zen et al., licensed under
CC BY 4.0
. It was used as clean-speech training data for the released checkpoint.
VCTK Corpus: the
VCTK dataset
from the Centre for Speech Technology Research, University of Edinburgh, licensed under
CC BY 4.0
. It was used as clean-speech training data for the released checkpoint.
DNS5 Challenge noise resources: noise data prepared through the
DNS Challenge
and used to synthesize noisy training mixtures for the released checkpoint. For this release, the DNS5 noise resources draw on
AudioSet
material licensed under
CC BY 4.0
, selected
Freesound
files licensed under
CC0 1.0
, and
DEMAND
environmental recordings licensed under
CC BY-SA 3.0
.
OpenSLR26 and OpenSLR28:
OpenSLR26
and
OpenSLR28
room impulse response resources, both licensed under Apache 2.0, were used to add reverberation during training.
The performance of the released version compared to the paper's results:
Model
DNSMOS
UTMOS
SBS
LPS
SpkSim
WER (%)
Vocoder-L24 (paper)
3.23
3.40
0.94
0.97
0.65
2.86
Vocoder-L24 (released)
3.29
3.30
0.94
0.96
0.59
3.46
DeWavLM (paper)
3.26
3.42
0.88
0.93
0.57
7.62
DeWavLM (released)
3.31
3.39
0.88
0.93
0.52
7.25
PASE (paper)
3.12
3.09
0.90
0.93
0.80
7.49
PASE (released)
3.08
3.21
0.91
0.94
0.80
6.76
It can be seen that the released version achieves performance very close to that of the paper's results on our simulated test set.
Overall, PASE achieves:
Lowest WER among evaluated generative and discriminative baselines
Highest speaker similarity (SpkSim)
Strong perceptual quality with low hallucination rates
Consistent performance across noisy and reverberant conditions
Bias, Risks, and Limitations
Model trained primarily on English speech; performance may degrade for other languages.
Very strong noise or mismatched reverberation conditions can introduce artifacts.
Speaker characteristics are preserved but not guaranteed perfectly.
Recommendations
Evaluate outputs for your specific use case. Avoid deployments where misunderstanding enhanced speech could have safety or legal consequences.
Citation
If you use PASE in your research, please cite:
@article{PASE,
title={{PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement}},
volume={40},
DOI={10.1609/aaai.v40i39.40562},
number={39},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
author={Rong, Xiaobin and Hu, Qinwen and Yesilbursa, Mansur and Wojcicki, Kamil and Lu, Jing},
year={2026},
month={Mar.},
pages={32826-32834}
}
pase huggingface.co is an AI model on huggingface.co that provides pase's model effect (), which can be used instantly with this cisco-ai pase model. huggingface.co supports a free trial of the pase model, and also provides paid use of the pase. Support call pase model through api, including Node.js, Python, http.
pase huggingface.co is an online trial and call api platform, which integrates pase's modeling effects, including api services, and provides a free online trial of pase, you can try pase online for free by clicking the link below.
pase is an open source model from GitHub that offers a free installation service, and any user can find pase on GitHub to install. At the same time, huggingface.co provides the effect of pase install, users can directly use pase installed effect in huggingface.co for debugging and trial. It also supports api for free installation.