Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
Lijie Liu
*
,
Tianxiang Ma
*
,
Bingchuan Li
*
,
Zhuowei Chen
*
,
Jiawei Liu
,
Gen Li
,
Siyu Zhou
,
Qian He
,
Xinglong Wu
*
Equal contribution,
†
Project lead
Intelligent Creation Team, ByteDance
🔥 Latest News!
Apr 20, 2025: 👋 Phantom-Wan is coming! We adapted the Phantom framework into the
Wan2.1
video generation model. The inference codes and checkpoint have been released.
📑 Todo List
Inference codes and Checkpoint of Phantom-Wan 1.3B
Checkpoint of Phantom-Wan 14B
Training codes of Phantom-Wan
📖 Overview
Phantom is a unified video generation framework for single and multi-subject references, built on existing text-to-video and image-to-video architectures. It achieves cross-modal alignment using text-image-video triplet data by redesigning the joint text-image injection model. Additionally, it emphasizes subject consistency in human generation while enhancing ID-preserving video generation.
⚡️ Quickstart
Installation
Clone the repo:
git clone https://github.com/Phantom-video/Phantom.git
cd Phantom
Changing
--ref_image
can achieve single reference Subject-to-Video generation or multi-reference Subject-to-Video generation. The number of reference images should be within 4.
To achieve the best generation results, we recommend that you describe the visual content of the reference image as accurately as possible when writing
--prompt
. For example, "examples/ref1.png" can be described as "a toy camera in yellow and red with blue buttons".
When the generated video is unsatisfactory, the most straightforward solution is to try changing the
--base_seed
and modifying the description in the
--prompt
.
For inferencing examples, please refer to "infer.sh". You will get the following generated results:
🆚 Comparative Results
Identity Preserving Video Generation
.
Single Reference Subject-to-Video Generation
.
Multi-Reference Subject-to-Video Generation
.
Acknowledgements
We would like to express our gratitude to the SEED team for their support. Special thanks to Lu Jiang, Haoyuan Guo, Zhibei Ma, and Sen Wang for their assistance with the model and data. In addition, we are also very grateful to Siying Chen, Qingyang Li, and Wei Han for their help with the evaluation.
BibTeX
@article{liu2025phantom,
title={Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment},
author={Liu, Lijie and Ma, Tianxaing and Li, Bingchuan and Chen, Zhuowei and Liu, Jiawei and He, Qian and Wu, Xinglong},
journal={arXiv preprint arXiv:2502.11079},
year={2025}
}
Runs of bytedance-research Phantom on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About Phantom huggingface.co Model
Phantom huggingface.co is an AI model on huggingface.co that provides Phantom's model effect (), which can be used instantly with this bytedance-research Phantom model. huggingface.co supports a free trial of the Phantom model, and also provides paid use of the Phantom. Support call Phantom model through api, including Node.js, Python, http.
Phantom huggingface.co is an online trial and call api platform, which integrates Phantom's modeling effects, including api services, and provides a free online trial of Phantom, you can try Phantom online for free by clicking the link below.
bytedance-research Phantom online free url in huggingface.co:
Phantom is an open source model from GitHub that offers a free installation service, and any user can find Phantom on GitHub to install. At the same time, huggingface.co provides the effect of Phantom install, users can directly use Phantom installed effect in huggingface.co for debugging and trial. It also supports api for free installation.