AudioLDM is a latent text-to-audio diffusion model capable of generating realistic audio samples given any text input. It is available in the 🧨 Diffusers library from v0.15.0 onwards.
Inspired by
Stable Diffusion
, AudioLDM
is a text-to-audio
latent diffusion model (LDM)
that learns continuous audio representations from
CLAP
latents. AudioLDM takes a text prompt as input and predicts the corresponding audio. It can generate text-conditional
sound effects, human speech and music.
Checkpoint Details
This is the original,
small
version of the AudioLDM model, also referred to as
audioldm-s-full
. The four AudioLDM checkpoints are summarised in the table below:
For text-to-audio generation, the
AudioLDMPipeline
can be
used to load pre-trained weights and generate text-conditional audio outputs:
from diffusers import AudioLDMPipeline
import torch
repo_id = "cvssp/audioldm"
pipe = AudioLDMPipeline.from_pretrained(repo_id, torch_dtype=torch.float16)
pipe = pipe.to("cuda")
prompt = "Techno music with a strong, upbeat tempo and high melodic riffs"
audio = pipe(prompt, num_inference_steps=10, audio_length_in_s=5.0).audios[0]
The resulting audio output can be saved as a .wav file:
Or displayed in a Jupyter Notebook / Google Colab:
from IPython.display import Audio
Audio(audio, rate=16000)
Tips
Prompts:
Descriptive prompt inputs work best: you can use adjectives to describe the sound (e.g. "high quality" or "clear") and make the prompt context specific (e.g., "water stream in a forest" instead of "stream").
It's best to use general terms like 'cat' or 'dog' instead of specific names or abstract objects that the model may not be familiar with.
Inference:
The
quality
of the predicted audio sample can be controlled by the
num_inference_steps
argument: higher steps give higher quality audio at the expense of slower inference.
The
length
of the predicted audio sample can be controlled by varying the
audio_length_in_s
argument.
Citation
BibTeX:
@article{liu2023audioldm,
title={AudioLDM: Text-to-Audio Generation with Latent Diffusion Models},
author={Liu, Haohe and Chen, Zehua and Yuan, Yi and Mei, Xinhao and Liu, Xubo and Mandic, Danilo and Wang, Wenwu and Plumbley, Mark D},
journal={arXiv preprint arXiv:2301.12503},
year={2023}
}
Runs of cvssp audioldm on huggingface.co
212
Total runs
2
24-hour runs
5
3-day runs
16
7-day runs
-126
30-day runs
More Information About audioldm huggingface.co Model
audioldm huggingface.co is an AI model on huggingface.co that provides audioldm's model effect (), which can be used instantly with this cvssp audioldm model. huggingface.co supports a free trial of the audioldm model, and also provides paid use of the audioldm. Support call audioldm model through api, including Node.js, Python, http.
audioldm huggingface.co is an online trial and call api platform, which integrates audioldm's modeling effects, including api services, and provides a free online trial of audioldm, you can try audioldm online for free by clicking the link below.
audioldm is an open source model from GitHub that offers a free installation service, and any user can find audioldm on GitHub to install. At the same time, huggingface.co provides the effect of audioldm install, users can directly use audioldm installed effect in huggingface.co for debugging and trial. It also supports api for free installation.