KVAE-Audio
is a continuous, full-band (48 kHz) audio autoencoder. It compresses raw waveforms into compact continuous latents and reconstructs them with high fidelity across speech, music, and general sound. The model is designed not only for faithful reconstruction, but as a
latent space for generative models
— in our internal text-to-audio pipeline, swapping the autoencoder for KVAE-Audio improves generation quality under a fixed generator.
Evaluation results
Evaluation of latent space qualities for generation
Generative quality is established under a
fixed generator
— same DiT architecture, training data, and number of steps — varying only the autoencoder. We report objective generation metrics and blind human side-by-side below.
KVAE-Audio huggingface.co is an AI model on huggingface.co that provides KVAE-Audio's model effect (), which can be used instantly with this kandinskylab KVAE-Audio model. huggingface.co supports a free trial of the KVAE-Audio model, and also provides paid use of the KVAE-Audio. Support call KVAE-Audio model through api, including Node.js, Python, http.
KVAE-Audio huggingface.co is an online trial and call api platform, which integrates KVAE-Audio's modeling effects, including api services, and provides a free online trial of KVAE-Audio, you can try KVAE-Audio online for free by clicking the link below.
kandinskylab KVAE-Audio online free url in huggingface.co:
KVAE-Audio is an open source model from GitHub that offers a free installation service, and any user can find KVAE-Audio on GitHub to install. At the same time, huggingface.co provides the effect of KVAE-Audio install, users can directly use KVAE-Audio installed effect in huggingface.co for debugging and trial. It also supports api for free installation.