Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
Bernini-Diffusers-v2
packages the full semantic-planning pipeline: a Qwen2.5-VL planner, Bernini planning weights, and Wan2.2 diffusion components in one self-contained diffusers-format directory.
Compared with the renderer-only Bernini-R releases, Bernini-Diffusers-v2 is recommended when you need stronger instruction following, multi-step semantic planning, and better handling of complex video generation or editing requests. Compared with the first Bernini-Diffusers release, v2 uses a training recipe that warms up the connector for thousands of steps before co-training, improving reference-guided video editing and OpenS2V performance.
🧾 Model card
Field
Description
Model type
Full video generation/editing pipeline with an MLLM-based semantic planner and a DiT-based renderer.
On video editing, Bernini reaches the first tier among leading closed-source commercial models in our internal arena evaluation based on blind human pairwise comparisons.
📦 Package layout
This release is a
self-contained diffusers-format directory
. Pass the downloaded
Bernini-Diffusers-v2
directory directly to
--config
.
git clone https://github.com/bytedance/Bernini.git bernini && cd bernini
pip install -r requirements.txt
# Open-VeOmni is required. Install it with --no-deps so it does not pull in a# different torch build and override the pinned torch==2.7.1+cu126:
pip install --no-deps git+https://github.com/ByteDance-Seed/[email protected]
Recommended environment:
Python
3.11.2
PyTorch
2.7.1+cu126
CUDA toolkit
12.6
GPU
Hopper GPUs (H100/H800/H200) are recommended for best performance
Load the model
Pass the downloaded directory directly as
--config
:
--use_pe
enhances the prompt through an OpenAI-compatible endpoint and is recommended for best generation quality.
export BERNINI_PE_API_KEY=... # or OPENAI_API_KEYexport BERNINI_PE_BASE_URL=... # or OPENAI_BASE_URLexport BERNINI_PE_MODEL=... # vision-capable chat model
@article{bernini,
title = {Bernini: Latent Semantic Planning for Video Diffusion},
author = {Chenchen Liu and Junyi Chen and Lei Li and Lu Chi and Mingzhen Sun and Zhuoying Li and Yi Fu and Ruoyu Guo and Yiheng Wu and Ge Bai and Zehuan Yuan},
journal = {arXiv preprint arXiv:2605.22344},
year = {2026}
}
🙏 Acknowledgements
Bernini builds on several outstanding open-source projects:
Bernini-Diffusers-v2 huggingface.co is an AI model on huggingface.co that provides Bernini-Diffusers-v2's model effect (), which can be used instantly with this ByteDance Bernini-Diffusers-v2 model. huggingface.co supports a free trial of the Bernini-Diffusers-v2 model, and also provides paid use of the Bernini-Diffusers-v2. Support call Bernini-Diffusers-v2 model through api, including Node.js, Python, http.
Bernini-Diffusers-v2 huggingface.co is an online trial and call api platform, which integrates Bernini-Diffusers-v2's modeling effects, including api services, and provides a free online trial of Bernini-Diffusers-v2, you can try Bernini-Diffusers-v2 online for free by clicking the link below.
ByteDance Bernini-Diffusers-v2 online free url in huggingface.co:
Bernini-Diffusers-v2 is an open source model from GitHub that offers a free installation service, and any user can find Bernini-Diffusers-v2 on GitHub to install. At the same time, huggingface.co provides the effect of Bernini-Diffusers-v2 install, users can directly use Bernini-Diffusers-v2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Bernini-Diffusers-v2 install url in huggingface.co: