January 26, 2026
: 🎉
HunyuanImage-3.0-Instruct
- Release of
Instruct (with reasoning)
for intelligent prompt enhancement and
Image-to-Image
generation for creative editing.
HunyuanImage-3.0
is a groundbreaking native multimodal model that unifies multimodal understanding and generation within an autoregressive framework. Our text-to-image and image-to-image model achieves performance
comparable to or surpassing
leading closed-source models.
✨ Key Features
🧠
Unified Multimodal Architecture:
Moving beyond the prevalent DiT-based architectures, HunyuanImage-3.0 employs a unified autoregressive framework. This design enables a more direct and integrated modeling of text and image modalities, leading to surprisingly effective and contextually rich image generation.
🏆
The Largest Image Generation MoE Model:
This is the largest open-source image generation Mixture of Experts (MoE) model to date. It features 64 experts and a total of 80 billion parameters, with 13 billion activated per token, significantly enhancing its capacity and performance.
🎨
Superior Image Generation Performance:
Through rigorous dataset curation and advanced reinforcement learning post-training, we've achieved an optimal balance between semantic accuracy and visual excellence. The model demonstrates exceptional prompt adherence while delivering photorealistic imagery with stunning aesthetic quality and fine-grained details.
💭
Intelligent Image Understanding and World-Knowledge Reasoning:
The unified multimodal architecture endows HunyuanImage-3.0 with powerful reasoning capabilities. It under stands user's input image, and leverages its extensive world knowledge to intelligently interpret user intent, automatically elaborating on sparse prompts with contextually appropriate details to produce superior, more complete visual outputs.
🚀 Usage
📦 Environment Setup
🐍
Python:
3.12+ (recommended and tested)
⚡
CUDA:
12.8
📥 Install Dependencies
# 1. First install PyTorch (CUDA 12.8 Version)
pip install torch==2.8.0 torchvision==0.23.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
# 2. Install tencentcloud-sdk for Prompt Enhancement (PE) only for HunyuanImage-3.0 not HunyuanImage-3.0-Instruct
pip install -i https://mirrors.tencent.com/pypi/simple/ --upgrade tencentcloud-sdk-python
# 3. Then install other dependencies
pip install -r requirements.txt
For
up to 3x faster inference
, install these optimizations:
# FlashInfer for optimized moe inference. v0.5.0 is tested.
pip install flashinfer-python==0.5.0
💡
Installation Tips:
It is critical that the CUDA version used by PyTorch matches the system's CUDA version.
FlashInfer relies on this compatibility when compiling kernels at runtime.
GCC version >=9 is recommended for compiling FlashAttention and FlashInfer.
⚡
Performance Tips:
These optimizations can significantly speed up your inference!
💡
Notation:
When FlashInfer is enabled, the first inference may be slower (about 10 minutes) due to kernel compilation. Subsequent inferences on the same machine will be much faster.
HunyuanImage-3.0-Instruct (Instruction reasoning and Image-to-image generation, including editing and multi-image fusion)
🔥 Quick Start with Transformers
1️⃣ Download model weights
# Download from HuggingFace and rename the directory.# Notice that the directory name should not contain dots, which may cause issues when loading using Transformers.
hf download tencent/HunyuanImage-3.0-Instruct --local-dir ./HunyuanImage-3-Instruct
2️⃣ Run with Transformers
from transformers import AutoModelForCausalLM
# Load the model
model_id = "./HunyuanImage-3-Instruct"# Currently we can not load the model using HF model_id `tencent/HunyuanImage-3.0-Instruct` directly # due to the dot in the name.
kwargs = dict(
attn_implementation="sdpa",
trust_remote_code=True,
torch_dtype="auto",
device_map="auto",
moe_impl="eager", # Use "flashinfer" if FlashInfer is installed
moe_drop_tokens=True,
)
model = AutoModelForCausalLM.from_pretrained(model_id, **kwargs)
model.load_tokenizer(model_id)
# Image-to-Image generation (TI2I)
prompt = "基于图一的logo,参考图二中冰箱贴的材质,制作一个新的冰箱贴"
input_img1 = "./assets/demo_instruct_imgs/input_1_0.png"
input_img2 = "./assets/demo_instruct_imgs/input_1_1.png"
imgs_input = [input_img1, input_img2]
cot_text, samples = model.generate_image(
prompt=prompt,
image=imgs_input,
seed=42,
image_size="auto",
use_system_prompt="en_unified",
bot_task="think_recaption", # Use "think_recaption" for reasoning and enhancement
infer_align_image_size=True, # Align output image size to input image size
diff_infer_steps=50,
verbose=2
)
# Save the generated image
samples[0].save("image_edit.png")
🏠 Local Installation & Usage
1️⃣ Clone the Repository
git clone https://github.com/Tencent-Hunyuan/HunyuanImage-3.0.git
cd HunyuanImage-3.0/
2️⃣ Download Model Weights
# Download from HuggingFace
hf download tencent/HunyuanImage-3.0-Instruct --local-dir ./HunyuanImage-3-Instruct
Custom system prompt. Used when
--use-system-prompt
is
custom
None
--bot-task
Task type.
image
for direct generation;
auto
for text;
recaption
for re-write->image;
think_recaption
for think->re-write->image
think_recaption
--save
Image save path
image.png
--verbose
Verbose level
2
--reproduce
Whether to reproduce the results
True
--infer-align-image-size
Whether to align the target image size to the src image size
True
--max_new_tokens
Maximum number of new tokens to generate
2048
--use-taylor-cache
Use Taylor Cache when sampling
False
5️⃣ For fewer Sampling Steps
We recommend using the model
HunyuanImage-3.0-Instruct-Distil
with
--diff-infer-steps 8
, while keeping all other recommended parameter values
unchanged
.
# Download HunyuanImage-3.0-Instruct-Distil from HuggingFace
hf download tencent/HunyuanImage-3.0-Instruct-Distil --local-dir ./HunyuanImage-3-Instruct-Distil
# Run the demo with 8 steps to samplesexport MODEL_PATH="./HunyuanImage-3-Instruct-Distil"
bash run_demo_instruct_Distil.sh
Previous Version (Pure Text-to-Image)
HunyuanImage-3.0 (Text-to-image)
🔥 Quick Start with Transformers
1️⃣ Download model weights
# Download from HuggingFace and rename the directory.# Notice that the directory name should not contain dots, which may cause issues when loading using Transformers.
hf download tencent/HunyuanImage-3.0 --local-dir ./HunyuanImage-3
2️⃣ Run with Transformers
from transformers import AutoModelForCausalLM
# Load the model
model_id = "./HunyuanImage-3"# Currently we can not load the model using HF model_id `tencent/HunyuanImage-3.0` directly # due to the dot in the name.
kwargs = dict(
attn_implementation="sdpa", # Use "flash_attention_2" if FlashAttention is installed
trust_remote_code=True,
torch_dtype="auto",
device_map="auto",
moe_impl="eager", # Use "flashinfer" if FlashInfer is installed
)
model = AutoModelForCausalLM.from_pretrained(model_id, **kwargs)
model.load_tokenizer(model_id)
# generate the image
prompt = "A brown and white dog is running on the grass"
image = model.generate_image(prompt=prompt, stream=True)
image.save("image.png")
🏠 Local Installation & Usage
1️⃣ Clone the Repository
git clone https://github.com/Tencent-Hunyuan/HunyuanImage-3.0.git
cd HunyuanImage-3.0/
2️⃣ Download Model Weights
# Download from HuggingFace
hf download tencent/HunyuanImage-3.0 --local-dir ./HunyuanImage-3
3️⃣ Run the Demo
The Pretrain Checkpoint does not automatically rewrite or enhance input prompts, for optimal results currently, we recommend community partners to use deepseek to rewrite the prompts. You can go to
Tencent Cloud
to apply for an API Key.
# Without PEexport MODEL_PATH="./HunyuanImage-3"
python3 run_image_gen.py \
--model-id $MODEL_PATH \
--verbose 1 \
--prompt "A brown and white dog is running on the grass" \
--bot-task image \
--image-size "1024x1024" \
--save ./image.png \
--moe-impl flashinfer
# With PEexport DEEPSEEK_KEY_ID="your_deepseek_key_id"export DEEPSEEK_KEY_SECRET="your_deepseek_key_secret"export MODEL_PATH="./HunyuanImage-3"
python3 run_image_gen.py \
--model-id $MODEL_PATH \
--verbose 1 \
--prompt "A brown and white dog is running on the grass" \
--bot-task image \
--image-size "1024x1024" \
--save ./image.png \
--moe-impl flashinfer \
--rewrite 1
4️⃣ Command Line Arguments
Arguments
Description
Recommended
--prompt
Input prompt
(Required)
--model-id
Model path
(Required)
--attn-impl
Attention implementation. Either
sdpa
or
flash_attention_2
.
sdpa
--moe-impl
MoE implementation. Either
eager
or
flashinfer
flashinfer
--seed
Random seed for image generation
None
--diff-infer-steps
Diffusion infer steps
50
--image-size
Image resolution. Can be
auto
, like
1280x768
or
16:9
auto
--save
Image save path.
image.png
--verbose
Verbose level. 0: No log; 1: log inference information.
0
--rewrite
Whether to enable rewriting
1
🎨 Interactive Gradio Demo
Launch an interactive web interface for easy text-to-image generation.
1️⃣ Install Gradio
pip install gradio>=4.21.0
2️⃣ Configure Environment
# Set your model pathexport MODEL_ID="path/to/your/model"# Optional: Configure GPU usage (default: 0,1,2,3)export GPUS="0,1,2,3"# Optional: Configure host and port (default: 0.0.0.0:443)export HOST="0.0.0.0"export PORT="443"
3️⃣ Launch the Web Interface
Basic Launch:
sh run_app.sh
With Performance Optimizations:
# Use both optimizations for maximum performance
sh run_app.sh --moe-impl flashinfer --attn-impl flash_attention_2
4️⃣ Access the Interface
🌐
Web Interface:
Open your browser and navigate to
http://localhost:443
(or your configured port)
Install performance extras (FlashAttention, FlashInfer) for faster inference.
Multi‑GPU inference is recommended for the Base model.
📊 Evaluation
Evaluation of HunyuanImage-3.0-Instruct
👥
GSB (Human Evaluation)
We adopted the GSB (Good/Same/Bad) evaluation method commonly used to assess the relative performance between two models from an overall image perception perspective. In total, we utilized 1,000+ single- and multi-images editing cases, generating an equal number of image samples for all compared models in a single run. For a fair comparison, we conducted inference only once for each prompt, avoiding any cherry-picking of results. When comparing with the baseline methods, we maintained the default settings for all selected models. The evaluation was performed by more than 100 professional evaluators.
Evaluation of HunyuanImage-3.0 (Text-to-Image)
🤖
SSAE (Machine Evaluation)
SSAE (Structured Semantic Alignment Evaluation) is an intelligent evaluation metric for image-text alignment based on advanced multimodal large language models (MLLMs). We extracted 3500 key points across 12 categories, then used multimodal large language models to automatically evaluate and score by comparing the generated images with these key points based on the visual content of the images. Mean Image Accuracy represents the image-wise average score across all key points, while Global Accuracy directly calculates the average score across all key points.
👥
GSB (Human Evaluation)
We adopted the GSB (Good/Same/Bad) evaluation method commonly used to assess the relative performance between two models from an overall image perception perspective. In total, we utilized 1,000 text prompts, generating an equal number of image samples for all compared models in a single run. For a fair comparison, we conducted inference only once for each prompt, avoiding any cherry-picking of results. When comparing with the baseline methods, we maintained the default settings for all selected models. The evaluation was performed by more than 100 professional evaluators.
🖼️ Showcase
Our model can follow complex instructions to generate high‑quality, creative images.
For text-to-image showcases in HunyuanImage-3.0, click the following links:
HunyuanImage-3.0-Instruct demonstrates powerful capabilities in intelligent image generation and editing. The following showcases highlight its core features:
🧠
Intelligent Visual Understanding and Reasoning (CoT Think)
: The model performs structured thinking to analyze user's input image and prompt, expand user's intent and editing tasks into a stucture, comprehnsive instructions, and leading to a better image generation and editing performance.
breaking down complex prompts and editing tasks into detailed visual components including subject, composition, lighting, color palette, and style.
✏️
Prompt Self-Rewrite
: Automatically enhances sparse or vague prompts into professional-grade, detail-rich descriptions that capture the user's intent more accurately.
🎨
Text-to-Image (T2I)
: Generates high-quality images from text prompts with exceptional prompt adherence and photorealistic quality.
🖼️
Image-to-Image (TI2I)
: Supports creative image editing, including adding elements, removing objects, modifying styles, and seamless background replacement while preserving key visual elements.
🔀
Multi-Image Fusion
: Intelligently combines multiple reference images (up to 3 inputs) to create coherent composite images that integrate visual elements from different sources.
Showcase 1: Detailed Thought and Reasoning Process
Showcase 2: Creative T2I Generation with Complex Scene Understanding
Prompt: 3D 毛绒质感拟人化马,暖棕浅棕肌理,穿藏蓝西装、白衬衫,戴深棕手套;疲惫带期待,坐于电脑前,旁置印 "HAPPY AGAIN" 的马克杯。橙红渐变背景,配超大号藏蓝粗体 "马上下班",叠加米黄 "Happy New Year" 并标 "(2026)"。橙红为主,藏蓝米黄撞色,毛绒温暖柔和。
Showcase 3: Precise Image Editing with Element Preservation
Showcase 4: Style Transformation with Thematic Enhancement
Showcase 5: Advanced Style Transfer and Product Mockup Generation
Showcase 6: Multi-Image Fusion and Creative Composition
📚 Citation
If you find HunyuanImage-3.0 useful in your research, please cite our work:
@article{cao2025hunyuanimage,
title={HunyuanImage 3.0 Technical Report},
author={Cao, Siyu and Chen, Hangting and Chen, Peng and Cheng, Yiji and Cui, Yutao and Deng, Xinchi and Dong, Ying and Gong, Kipper and Gu, Tianpeng and Gu, Xiusen and others},
journal={arXiv preprint arXiv:2509.23951},
year={2025}
}
🙏 Acknowledgements
We extend our heartfelt gratitude to the following open-source projects and communities for their invaluable contributions:
HunyuanImage-3-Instruct-verbatim-flashpack huggingface.co is an AI model on huggingface.co that provides HunyuanImage-3-Instruct-verbatim-flashpack's model effect (), which can be used instantly with this fal HunyuanImage-3-Instruct-verbatim-flashpack model. huggingface.co supports a free trial of the HunyuanImage-3-Instruct-verbatim-flashpack model, and also provides paid use of the HunyuanImage-3-Instruct-verbatim-flashpack. Support call HunyuanImage-3-Instruct-verbatim-flashpack model through api, including Node.js, Python, http.
HunyuanImage-3-Instruct-verbatim-flashpack huggingface.co is an online trial and call api platform, which integrates HunyuanImage-3-Instruct-verbatim-flashpack's modeling effects, including api services, and provides a free online trial of HunyuanImage-3-Instruct-verbatim-flashpack, you can try HunyuanImage-3-Instruct-verbatim-flashpack online for free by clicking the link below.
fal HunyuanImage-3-Instruct-verbatim-flashpack online free url in huggingface.co:
HunyuanImage-3-Instruct-verbatim-flashpack is an open source model from GitHub that offers a free installation service, and any user can find HunyuanImage-3-Instruct-verbatim-flashpack on GitHub to install. At the same time, huggingface.co provides the effect of HunyuanImage-3-Instruct-verbatim-flashpack install, users can directly use HunyuanImage-3-Instruct-verbatim-flashpack installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
HunyuanImage-3-Instruct-verbatim-flashpack install url in huggingface.co: