CogVideoX is an open-source video generation model originating
from
Qingying
. The table below presents information related to the video
generation models we offer in this version.
Model Name
CogVideoX-2B
CogVideoX-5B
CogVideoX-5B-I2V (This Repository)
Model Description
Entry-level model, balancing compatibility. Low cost for running and secondary development.
Larger model with higher video generation quality and better visual effects.
CogVideoX-5B image-to-video version.
Inference Precision
FP16*(recommended)
, BF16, FP32, FP8*, INT8, not supported: INT4
BF16 (recommended)
, FP16, FP32, FP8*, INT8, not supported: INT4
Single GPU Memory Usage
SAT
FP16: 18GB
diffusers FP16: from 4GB*
diffusers INT8 (torchao): from 3.6GB*
SAT
BF16: 26GB
diffusers BF16: from 5GB*
diffusers INT8 (torchao): from 4.4GB*
Multi-GPU Inference Memory Usage
FP16: 10GB* using diffusers
BF16: 15GB* using diffusers
Inference Speed
(Step = 50, FP/BF16)
Single A100: ~90 seconds
Single H100: ~45 seconds
Single A100: ~180 seconds
Single H100: ~90 seconds
720 x 480, no support for other resolutions (including fine-tuning)
Position Embedding
3d_sincos_pos_embed
3d_rope_pos_embed
3d_rope_pos_embed + learnable_pos_embed
Data Explanation
While testing using the diffusers library, all optimizations included in the diffusers library were enabled. This
scheme has not been tested for actual memory usage on devices outside of
NVIDIA A100 / H100
architectures.
Generally, this scheme can be adapted to all
NVIDIA Ampere architecture
and above devices. If optimizations are
disabled, memory consumption will multiply, with peak memory usage being about 3 times the value in the table.
However, speed will increase by about 3-4 times. You can selectively disable some optimizations, including:
For multi-GPU inference, the
enable_sequential_cpu_offload()
optimization needs to be disabled.
Using INT8 models will slow down inference, which is done to accommodate lower-memory GPUs while maintaining minimal
video quality loss, though inference speed will significantly decrease.
The CogVideoX-2B model was trained in
FP16
precision, and all CogVideoX-5B models were trained in
BF16
precision.
We recommend using the precision in which the model was trained for inference.
PytorchAO
and
Optimum-quanto
can be
used to quantize the text encoder, transformer, and VAE modules to reduce the memory requirements of CogVideoX. This
allows the model to run on free T4 Colabs or GPUs with smaller memory! Also, note that TorchAO quantization is fully
compatible with
torch.compile
, which can significantly improve inference speed. FP8 precision must be used on
devices with NVIDIA H100 and above, requiring source installation of
torch
,
torchao
,
diffusers
, and
accelerate
Python packages. CUDA 12.4 is recommended.
The inference speed tests also used the above memory optimization scheme. Without memory optimization, inference speed
increases by about 10%. Only the
diffusers
version of the model supports quantization.
The model only supports English input; other languages can be translated into English for use via large model
refinement.
The memory usage of model fine-tuning is tested in an
8 * H100
environment, and the program automatically
uses
Zero 2
optimization. If a specific number of GPUs is marked in the table, that number or more GPUs must be used
for fine-tuning.
Reminders
Use
SAT
for inference and fine-tuning SAT version models. Feel free
to visit our GitHub for more details.
Getting Started Quickly 🤗
This model supports deployment using the Hugging Face diffusers library. You can follow the steps below to get started.
We recommend that you visit our
GitHub
to check out prompt optimization and
conversion to get a better experience.
import torch
from diffusers import CogVideoXImageToVideoPipeline
from diffusers.utils import export_to_video, load_image
prompt = "A little girl is riding a bicycle at high speed. Focused, detailed, realistic."
image = load_image(image="input.jpg")
pipe = CogVideoXImageToVideoPipeline.from_pretrained(
"THUDM/CogVideoX-5b-I2V",
torch_dtype=torch.bfloat16
)
pipe.enable_sequential_cpu_offload()
pipe.vae.enable_tiling()
pipe.vae.enable_slicing()
video = pipe(
prompt=prompt,
image=image,
num_videos_per_prompt=1,
num_inference_steps=50,
num_frames=49,
guidance_scale=6,
generator=torch.Generator(device="cuda").manual_seed(42),
).frames[0]
export_to_video(video, "output.mp4", fps=8)
Quantized Inference
PytorchAO
and
Optimum-quanto
can be
used to quantize the text encoder, transformer, and VAE modules to reduce CogVideoX's memory requirements. This allows
the model to run on free T4 Colab or GPUs with lower VRAM! Also, note that TorchAO quantization is fully compatible
with
torch.compile
, which can significantly accelerate inference.
# To get started, PytorchAO needs to be installed from the GitHub source and PyTorch Nightly.
# Source and nightly installation is only required until the next release.
import torch
from diffusers import AutoencoderKLCogVideoX, CogVideoXTransformer3DModel, CogVideoXImageToVideoPipeline
from diffusers.utils import export_to_video, load_image
from transformers import T5EncoderModel
from torchao.quantization import quantize_, int8_weight_only
quantization = int8_weight_only
text_encoder = T5EncoderModel.from_pretrained("THUDM/CogVideoX-5b-I2V", subfolder="text_encoder", torch_dtype=torch.bfloat16)
quantize_(text_encoder, quantization())
transformer = CogVideoXTransformer3DModel.from_pretrained("THUDM/CogVideoX-5b-I2V",subfolder="transformer", torch_dtype=torch.bfloat16)
quantize_(transformer, quantization())
vae = AutoencoderKLCogVideoX.from_pretrained("THUDM/CogVideoX-5b-I2V", subfolder="vae", torch_dtype=torch.bfloat16)
quantize_(vae, quantization())
# Create pipeline and run inference
pipe = CogVideoXImageToVideoPipeline.from_pretrained(
"THUDM/CogVideoX-5b-I2V",
text_encoder=text_encoder,
transformer=transformer,
vae=vae,
torch_dtype=torch.bfloat16,
)
pipe.enable_model_cpu_offload()
pipe.vae.enable_tiling()
pipe.vae.enable_slicing()
prompt = "A little girl is riding a bicycle at high speed. Focused, detailed, realistic."
image = load_image(image="input.jpg")
video = pipe(
prompt=prompt,
image=image,
num_videos_per_prompt=1,
num_inference_steps=50,
num_frames=49,
guidance_scale=6,
generator=torch.Generator(device="cuda").manual_seed(42),
).frames[0]
export_to_video(video, "output.mp4", fps=8)
Additionally, these models can be serialized and stored using PytorchAO in quantized data types to save disk space. You
can find examples and benchmarks at the following links:
@article{yang2024cogvideox,
title={CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer},
author={Yang, Zhuoyi and Teng, Jiayan and Zheng, Wendi and Ding, Ming and Huang, Shiyu and Xu, Jiazheng and Yang, Yuanming and Hong, Wenyi and Zhang, Xiaohan and Feng, Guanyu and others},
journal={arXiv preprint arXiv:2408.06072},
year={2024}
}
Runs of THUDM CogVideoX-5b-I2V on huggingface.co
54.2K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About CogVideoX-5b-I2V huggingface.co Model
CogVideoX-5b-I2V huggingface.co is an AI model on huggingface.co that provides CogVideoX-5b-I2V's model effect (), which can be used instantly with this THUDM CogVideoX-5b-I2V model. huggingface.co supports a free trial of the CogVideoX-5b-I2V model, and also provides paid use of the CogVideoX-5b-I2V. Support call CogVideoX-5b-I2V model through api, including Node.js, Python, http.
CogVideoX-5b-I2V huggingface.co is an online trial and call api platform, which integrates CogVideoX-5b-I2V's modeling effects, including api services, and provides a free online trial of CogVideoX-5b-I2V, you can try CogVideoX-5b-I2V online for free by clicking the link below.
THUDM CogVideoX-5b-I2V online free url in huggingface.co:
CogVideoX-5b-I2V is an open source model from GitHub that offers a free installation service, and any user can find CogVideoX-5b-I2V on GitHub to install. At the same time, huggingface.co provides the effect of CogVideoX-5b-I2V install, users can directly use CogVideoX-5b-I2V installed effect in huggingface.co for debugging and trial. It also supports api for free installation.