Each line contains an utterance ID and an audio path. Audio is processed as
mono 24 kHz input and prepared as training segments by the dataset pipeline.
Adjust training settings in
configs/config_Sphere_VAE.yaml
. Checkpoints and
TensorBoard logs are written to the configured experiment directory. For
multi-GPU training, configure Accelerate and set the visible devices as needed.
--input
accepts an audio file or directory. Supported extensions are WAV,
FLAC, MP3, OGG, and M4A. Use
--config
to select another configuration.
Inference uses posterior sampling, so repeated runs can produce different
reconstructions. Output is trimmed to the input length.
Layout
models/model_Sphere_VAE.py Sphere_VAE model and builder
modules/ SEANet, Transformer, spherical distribution, utilities
train.py Training entrypoint
infer.py Audio reconstruction entrypoint
dataset.py SCP-listed audio dataset
configs/ Sphere_VAE configuration
discriminators/ STFT discriminator
losses/ Spectral and adversarial losses
utils/ Configuration, checkpoints, and compilation utilities
Citation
@misc{zhang2026spherevaehypersphericallatentautoencoders,
title={SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling},
author={Haoyu Zhang and Jingbin Hu and Hanke Xie and Qirui Zhan and Wenhao Li and Ziyu Zhang and Xiaming Ren and Yue Li and Xunyu Zhu and Zhipeng Chen and Lei Xie},
year={2026},
eprint={2609.09903},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2609.09903}
}
Acknowledgments
Parts of the code are adapted from Kyutai Mimi and Meta AudioCraft/EnCodec,
as indicated by the original notices retained in the source files. This
export does not assign a new license to that third-party code. The source
checkout did not include the root license files referenced by those notices;
applicable upstream license texts still need to be supplied before release.
Runs of ASLP-lab SphereVAE on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About SphereVAE huggingface.co Model
SphereVAE huggingface.co is an AI model on huggingface.co that provides SphereVAE's model effect (), which can be used instantly with this ASLP-lab SphereVAE model. huggingface.co supports a free trial of the SphereVAE model, and also provides paid use of the SphereVAE. Support call SphereVAE model through api, including Node.js, Python, http.
SphereVAE huggingface.co is an online trial and call api platform, which integrates SphereVAE's modeling effects, including api services, and provides a free online trial of SphereVAE, you can try SphereVAE online for free by clicking the link below.
ASLP-lab SphereVAE online free url in huggingface.co:
SphereVAE is an open source model from GitHub that offers a free installation service, and any user can find SphereVAE on GitHub to install. At the same time, huggingface.co provides the effect of SphereVAE install, users can directly use SphereVAE installed effect in huggingface.co for debugging and trial. It also supports api for free installation.