chenxwh / hunyuandit

A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

replicate.com
Total runs: 333
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Github
Model's Last Updated: May 24 2024

Introduction of hunyuandit

Model Details of hunyuandit

Readme

Hunyuan-DiT

A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

This repo contains PyTorch model definitions, pre-trained weights and inference/sampling code for our paper exploring Hunyuan-DiT. You can find more visualizations on our project page .

Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

DialogGen:Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation

Abstract

We present Hunyuan-DiT, a text-to-image diffusion transformer with fine-grained understanding of both English and Chinese. To construct Hunyuan-DiT, we carefully designed the transformer structure, text encoder, and positional encoding. We also build from scratch a whole data pipeline to update and evaluate data for iterative model optimization. For fine-grained language understanding, we train a Multimodal Large Language Model to refine the captions of the images. Finally, Hunyuan-DiT can perform multi-round multi-modal dialogue with users, generating and refining images according to the context. Through our carefully designed holistic human evaluation protocol with more than 50 professional human evaluators, Hunyuan-DiT sets a new state-of-the-art in Chinese-to-image generation compared with other open-source models.

🎉 Hunyuan-DiT Key Features
Chinese-English Bilingual DiT Architecture

Hunyuan-DiT is a diffusion model in the latent space, as depicted in figure below. Following the Latent Diffusion Model, we use a pre-trained Variational Autoencoder (VAE) to compress the images into low-dimensional latent spaces and train a diffusion model to learn the data distribution with diffusion models. Our diffusion model is parameterized with a transformer. To encode the text prompts, we leverage a combination of pre-trained bilingual (English and Chinese) CLIP and multilingual T5 encoder.

Multi-turn Text2Image Generation

Understanding natural language instructions and performing multi-turn interaction with users are important for a text-to-image system. It can help build a dynamic and iterative creation process that bring the user’s idea into reality step by step. In this section, we will detail how we empower Hunyuan-DiT with the ability to perform multi-round conversations and image generation. We train MLLM to understand the multi-round user dialogue and output the new text prompt for image generation.

📈 Comparisons

In order to comprehensively compare the generation capabilities of HunyuanDiT and other models, we constructed a 4-dimensional test set, including Text-Image Consistency, Excluding AI Artifacts, Subject Clarity, Aesthetic. More than 50 professional evaluators performs the evaluation.

🔗 BibTeX

If you find Hunyuan-DiT or DialogGen useful for your research and applications, please cite using this BibTeX:

@misc{li2024hunyuandit,
      title={Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding}, 
      author={Zhimin Li and Jianwei Zhang and Qin Lin and Jiangfeng Xiong and Yanxin Long and Xinchi Deng and Yingfang Zhang and Xingchao Liu and Minbin Huang and Zedong Xiao and Dayou Chen and Jiajun He and Jiahao Li and Wenyue Li and Chen Zhang and Rongwei Quan and Jianxiang Lu and Jiabin Huang and Xiaoyan Yuan and Xiaoxiao Zheng and Yixuan Li and Jihong Zhang and Chao Zhang and Meng Chen and Jie Liu and Zheng Fang and Weiyan Wang and Jinbao Xue and Yangyu Tao and Jianchen Zhu and Kai Liu and Sihuan Lin and Yifu Sun and Yun Li and Dongdong Wang and Mingtao Chen and Zhichao Hu and Xiao Xiao and Yan Chen and Yuhong Liu and Wei Liu and Di Wang and Yong Yang and Jie Jiang and Qinglin Lu},
      year={2024},
      eprint={2405.08748},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

@article{huang2024dialoggen,
  title={DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation},
  author={Huang, Minbin and Long, Yanxin and Deng, Xinchi and Chu, Ruihang and Xiong, Jiangfeng and Liang, Xiaodan and Cheng, Hong and Lu, Qinglin and Liu, Wei},
  journal={arXiv preprint arXiv:2403.08857},
  year={2024}
}

Pricing of hunyuandit replicate.com

Run time and cost

This model costs approximately $0.10 to run on Replicate, or 10 runs per $1, but this varies depending on your inputs. It is also open source and you can run it on your own computer with Docker .

This model runs on Nvidia A40 (Large) GPU hardware . Predictions typically complete within 143 seconds. The predict time for this model varies significantly based on the inputs.

Runs of chenxwh hunyuandit on replicate.com

333
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About hunyuandit replicate.com Model

hunyuandit replicate.com

hunyuandit replicate.com is an AI model on replicate.com that provides hunyuandit's model effect (A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding), which can be used instantly with this chenxwh hunyuandit model. replicate.com supports a free trial of the hunyuandit model, and also provides paid use of the hunyuandit. Support call hunyuandit model through api, including Node.js, Python, http.

hunyuandit replicate.com Url

https://replicate.com/chenxwh/hunyuandit

chenxwh hunyuandit online free

hunyuandit replicate.com is an online trial and call api platform, which integrates hunyuandit's modeling effects, including api services, and provides a free online trial of hunyuandit, you can try hunyuandit online for free by clicking the link below.

chenxwh hunyuandit online free url in replicate.com:

https://replicate.com/chenxwh/hunyuandit

hunyuandit install

hunyuandit is an open source model from GitHub that offers a free installation service, and any user can find hunyuandit on GitHub to install. At the same time, replicate.com provides the effect of hunyuandit install, users can directly use hunyuandit installed effect in replicate.com for debugging and trial. It also supports api for free installation.

hunyuandit install url in replicate.com:

https://replicate.com/chenxwh/hunyuandit

hunyuandit install url in github:

https://github.com/chenxwh/HunyuanDiT

Url of hunyuandit

Provider of hunyuandit replicate.com

Other API from chenxwh

replicate

Fast sdxl with higher quality

Total runs: 729.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:May 27 2024
replicate

CogVLM2: Visual Language Models for Image and Video Understanding

Total runs: 650.6K
Run Growth: 0
Growth Rate: 0.00%
Updated:September 25 2024
replicate

Depth estimation with faster inference speed, fewer parameters, and higher depth accuracy.

Total runs: 194.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:June 30 2024
replicate

Updated to OpenVoice v2: Versatile Instant Voice Cloning

Total runs: 55.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:May 18 2024
replicate

Audio-based Lip Synchronization for Talking Head Video

Total runs: 28.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 16 2024
replicate

Fast and High-Quality Text-to-video Generation

Total runs: 4.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:June 01 2024
replicate

OmniGen: Unified Image Generation

Total runs: 4.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 04 2024
replicate

Scalable Streaming Speech Synthesis with Large Language Models

Total runs: 3.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:December 26 2024
replicate

DiT-based video generation model for generating high-quality videos in real-time

Total runs: 2.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 24 2024
replicate

Convert LLM's coding to image generation

Total runs: 1.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:June 03 2024
replicate

Sharp Monocular Metric Depth in Less Than a Second

Total runs: 1.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:October 13 2024
replicate

Minimal and Universal Control for Diffusion Transformer - demo for Subject-driven generation

Total runs: 1.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 01 2025
replicate

Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Total runs: 1.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:December 29 2024
replicate

CogVLM2: Visual Language Models for Image and Video Understanding

Total runs: 573
Run Growth: 0
Growth Rate: 0.00%
Updated:September 25 2024
replicate

Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer

Total runs: 358
Run Growth: 0
Growth Rate: 0.00%
Updated:November 24 2024
replicate

Extended video synthesis model that generates 128 frames

Total runs: 203
Run Growth: 0
Growth Rate: 0.00%
Updated:July 01 2024
replicate

Depth Any Video with Scalable Synthetic Data

Total runs: 150
Run Growth: 0
Growth Rate: 0.00%
Updated:October 20 2024
replicate

High-Fidelity Single-Step Diffusion through Dynamic Adversarial Training

Total runs: 147
Run Growth: 0
Growth Rate: 0.00%
Updated:December 07 2024
replicate

Generating Consistent Long Depth Sequences for Open-world Videos

Total runs: 141
Run Growth: 0
Growth Rate: 0.00%
Updated:October 01 2024
replicate

One Diffusion to Generate Them All

Total runs: 135
Run Growth: 0
Growth Rate: 0.00%
Updated:December 31 2024
replicate

Diffusion-based Visual Foundation Model for High-quality Dense Prediction

Total runs: 131
Run Growth: 0
Growth Rate: 0.00%
Updated:October 07 2024
replicate

Efficient Visual Generation with Hybrid Autoregressive Transformer

Total runs: 121
Run Growth: 0
Growth Rate: 0.00%
Updated:October 19 2024
replicate

Minimal and Universal Control for Diffusion Transformer - demo for Spatially aligned control

Total runs: 96
Run Growth: 0
Growth Rate: 0.00%
Updated:January 01 2025
replicate

Image-to-Video Diffusion Models with An Expert Transformer

Total runs: 74
Run Growth: 0
Growth Rate: 0.00%
Updated:September 22 2024
replicate

Finer and Faster Text-to-Image Generation via Relay Diffusion

Total runs: 44
Run Growth: 0
Growth Rate: 0.00%
Updated:October 15 2024
replicate

Text-to-Video Diffusion Models with An Expert Transformer

Total runs: 42
Run Growth: 0
Growth Rate: 0.00%
Updated:September 21 2024
replicate

Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis

Total runs: 36
Run Growth: 0
Growth Rate: 0.00%
Updated:October 21 2024
replicate

Autoregressive Video Generation without Vector Quantization

Total runs: 32
Run Growth: 0
Growth Rate: 0.00%
Updated:December 27 2024
replicate

Emu3-Gen for image generation

Total runs: 27
Run Growth: 0
Growth Rate: 0.00%
Updated:September 30 2024
replicate

Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution

Total runs: 20
Run Growth: 0
Growth Rate: 0.00%
Updated:September 22 2024
replicate

Emu3-Chat for vision-language understanding

Total runs: 18
Run Growth: 0
Growth Rate: 0.00%
Updated:September 30 2024
replicate

Autoregressive Image Generation without Vector Quantization

Total runs: 14
Run Growth: 0
Growth Rate: 0.00%
Updated:December 27 2024
replicate

Let Vision Language Models Reason Step-by-Step

Total runs: 13
Run Growth: 0
Growth Rate: 0.00%
Updated:December 02 2024
replicate

Enhancing Video Model Post-Training through Data, Reward, and Conditional Guidance Design

Total runs: 10
Run Growth: 0
Growth Rate: 0.00%
Updated:October 14 2024
replicate

Text-to-Video Diffusion Models with An Expert Transformer

Total runs: 4
Run Growth: 0
Growth Rate: 0.00%
Updated:September 21 2024