vantagewithai / Ditto-GGUF

huggingface.co
Total runs: 316
24-hour runs: 9
7-day runs: 38
30-day runs: 102
Model's Last Updated: November 21 2025
video-to-video

Introduction of Ditto-GGUF

Model Details of Ditto-GGUF

GGUF Quantized versions of Ditto Models

Original model link: QingyanBai/Ditto_models

Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

This repository contains the Ditto framework and the Editto model, which are introduced in the paper Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset . Ditto provides a holistic approach to address the scarcity of high-quality training data for instruction-based video editing, enabling the creation of the Ditto-1M dataset and the training of the state-of-the-art Editto model.

Figure: Instruction-based video editing results produced by our proposed Ditto framework, which enables high-quality video editing through natural language instructions.

Abstract

Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this fundamental challenge. At its heart, Ditto features a novel data generation pipeline that fuses the creative diversity of a leading image editor with an in-context video generator, overcoming the limited scope of existing models. To make this process viable, our framework resolves the prohibitive cost-quality trade-off by employing an efficient, distilled model architecture augmented by a temporal enhancer, which simultaneously reduces computational overhead and improves temporal coherence. Finally, to achieve full scalability, this entire pipeline is driven by an intelligent agent that crafts diverse instructions and rigorously filters the output, ensuring quality control at scale. Using this framework, we invested over 12,000 GPU-days to build Ditto-1M, a new dataset of one million high-fidelity video editing examples. We trained our model, Editto, on Ditto-1M with a curriculum learning strategy. The results demonstrate superior instruction-following ability and establish a new state-of-the-art in instruction-based video editing.

Model Usage
1. Using with DiffSynth
Environment Setup
# Create conda environment (if you already have a DiffSynth conda environment, you can reuse it)
conda create -n ditto python=3.10
conda activate ditto
pip install -e .
Download Models

Download the base model and our models from Google Drive or Hugging Face :

# Download Wan-AI/Wan2.1-VACE-14B from Hugging Face to models/Wan-AI/
huggingface-cli download Wan-AI/Wan2.1-VACE-14B --local-dir models/Wan-AI/

# Download Ditto models
huggingface-cli download QingyanBai/Ditto_models --include="models/*" --local-dir ./
Usage

You can either use the provided script or run Python directly:

# Option 1: Use the provided script
bash infer.sh

# Option 2: Run Python directly
python inference/infer_ditto.py \
    --input_video /path/to/input_video.mp4 \
    --output_video /path/to/output_video.mp4 \
    --prompt "Editing instruction." \
    --lora_path /path/to/model.safetensors \
    --num_frames 73 \
    --device_id 0

Some test cases could be found at HF Dataset . You can also find some reference editing prompts in inference/example_prompts.txt .

2. Using with ComfyUI

Note: While ComfyUI runs faster with lower computational requirements (832×480x73 videos need 11G GPU memory and ~4min on A6000), please note that due to the use of quantized and distilled models, there may be some quality degradation.

Environment Setup

First, follow the ComfyUI installation guide to set up the base ComfyUI environment. We strongly recommend installing ComfyUI-Manager for easy custom node management:

# Install ComfyUI-Manager
cd ComfyUI/custom_nodes
git clone https://github.com/Comfy-Org/ComfyUI-Manager.git

After installing ComfyUI, you can either:

Option 1 (Recommended): Use ComfyUI-Manager to automatically install all required custom nodes with the function Install Missing Custom Nodes.

Option 2: Manually install the required custom nodes (you can refer to this page ):

Download Models

Download the required model weights from: Kijai/WanVideo_comfy to subfolders of models/ . Required files include:

Download our models from Google Drive or Hugging Face to diffusion_models/ (use VACE Module Select node for loading).

Usage

Use the workflow ditto_comfyui_workflow.json in this repo to get started. We provided some reference prompts in the note. Some test cases could be found at HF Dataset .

Note: If you want to test sim2real cases, you can try prompts like 'Turn it into the real domain'.

Citation

If you find this work useful, please consider citing our paper:

@article{bai2025ditto,
  title={Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset},
  author={Bai, Qingyan and Wang, Qiuyu and Ouyang, Hao and Yu, Yue and Wang, Hanlin and Wang, Wen and Cheng, Ka Leong and Ma, Shuailei and Zeng, Yanhong and Liu, Zichen and Xu, Yinghao and Shen, Yujun and Chen, Qifeng},
  journal={arXiv preprint arXiv:2510.15742},
  year={2025}
}
Acknowledgments

We thank Wan & VACE & Qwen-Image for providing the powerful foundation model, and QwenVL for the advanced visual understanding capabilities. We also thank DiffSynth-Studio serving as the codebase for this repository.

License

This project is licensed under the CC BY-NC-SA 4.0( Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License ).

The code is provided for academic research purposes only.

For any questions, please contact [email protected] .

Runs of vantagewithai Ditto-GGUF on huggingface.co

316
Total runs
9
24-hour runs
19
3-day runs
38
7-day runs
102
30-day runs

More Information About Ditto-GGUF huggingface.co Model

Ditto-GGUF huggingface.co

Ditto-GGUF huggingface.co is an AI model on huggingface.co that provides Ditto-GGUF's model effect (), which can be used instantly with this vantagewithai Ditto-GGUF model. huggingface.co supports a free trial of the Ditto-GGUF model, and also provides paid use of the Ditto-GGUF. Support call Ditto-GGUF model through api, including Node.js, Python, http.

vantagewithai Ditto-GGUF online free

Ditto-GGUF huggingface.co is an online trial and call api platform, which integrates Ditto-GGUF's modeling effects, including api services, and provides a free online trial of Ditto-GGUF, you can try Ditto-GGUF online for free by clicking the link below.

vantagewithai Ditto-GGUF online free url in huggingface.co:

https://huggingface.co/vantagewithai/Ditto-GGUF

Ditto-GGUF install

Ditto-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Ditto-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Ditto-GGUF install, users can directly use Ditto-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Ditto-GGUF install url in huggingface.co:

https://huggingface.co/vantagewithai/Ditto-GGUF

Url of Ditto-GGUF

Provider of Ditto-GGUF huggingface.co

vantagewithai
ORGANIZATIONS

Other API from vantagewithai