VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
[TL;DR]: VidMuse is a framework for generating high-fidelity music aligned with video content, utilizing Long-Short-Term modeling, and has been accepted to CVPR 2025.
sudo apt-get install ffmpeg
# Or if you are using Anaconda or Miniconda
conda install "ffmpeg<5" -c conda-forge
Run the following Python code:
from video_processor import VideoProcessor, merge_video_audio
from audiocraft.models import VidMuse
import scipy
# Path to the video
video_path = 'sample.mp4'# Initialize the video processor
processor = VideoProcessor()
# Process the video to obtain tensors and duration
local_video_tensor, global_video_tensor, duration = processor.process(video_path)
progress = True
USE_DIFFUSION = False# Load the pre-trained VidMuse model
MODEL = VidMuse.get_pretrained('HKUSTAudio/VidMuse')
# Set generation parameters for the model based on video duration
MODEL.set_generation_params(duration=duration)
try:
# Generate outputs using the model
outputs = MODEL.generate([local_video_tensor, global_video_tensor], progress=progress, return_tokens=USE_DIFFUSION)
except RuntimeError as e:
print(e)
# Detach outputs from the computation graph and convert to CPU float tensor
outputs = outputs.detach().cpu().float()
sampling_rate = 32000
output_wav_path = "vidmuse_sample.wav"# Write the output audio data to a WAV file
scipy.io.wavfile.write(output_wav_path, rate=sampling_rate, data=outputs[0, 0].numpy())
output_video_path = "vidmuse_sample.mp4"# Merge the original video with the generated music
merge_video_audio(video_path, output_wav_path, output_video_path)
Citation
If you find our work useful, please consider citing:
@article{tian2024vidmuse,
title={Vidmuse: A simple video-to-music generation framework with long-short-term modeling},
author={Tian, Zeyue and Liu, Zhaoyang and Yuan, Ruibin and Pan, Jiahao and Liu, Qifeng and Tan, Xu and Chen, Qifeng and Xue, Wei and Guo, Yike},
journal={arXiv preprint arXiv:2406.04321},
year={2024}
}
Runs of HKUSTAudio VidMuse on huggingface.co
1.3K
Total runs
0
24-hour runs
2
3-day runs
6
7-day runs
1.2K
30-day runs
More Information About VidMuse huggingface.co Model
VidMuse huggingface.co is an AI model on huggingface.co that provides VidMuse's model effect (), which can be used instantly with this HKUSTAudio VidMuse model. huggingface.co supports a free trial of the VidMuse model, and also provides paid use of the VidMuse. Support call VidMuse model through api, including Node.js, Python, http.
VidMuse huggingface.co is an online trial and call api platform, which integrates VidMuse's modeling effects, including api services, and provides a free online trial of VidMuse, you can try VidMuse online for free by clicking the link below.
HKUSTAudio VidMuse online free url in huggingface.co:
VidMuse is an open source model from GitHub that offers a free installation service, and any user can find VidMuse on GitHub to install. At the same time, huggingface.co provides the effect of VidMuse install, users can directly use VidMuse installed effect in huggingface.co for debugging and trial. It also supports api for free installation.