HunyuanImage-3.0
is a groundbreaking native multimodal model that unifies multimodal understanding and generation within an autoregressive framework. Our text-to-image module achieves performance
comparable to or surpassing
leading closed-source models.
โจ Key Features
๐ง
Unified Multimodal Architecture:
Moving beyond the prevalent DiT-based architectures, HunyuanImage-3.0 employs a unified autoregressive framework. This design enables a more direct and integrated modeling of text and image modalities, leading to surprisingly effective and contextually rich image generation.
๐
The Largest Image Generation MoE Model:
This is the largest open-source image generation Mixture of Experts (MoE) model to date. It features 64 experts and a total of 80 billion parameters, with 13 billion activated per token, significantly enhancing its capacity and performance.
๐จ
Superior Image Generation Performance:
Through rigorous dataset curation and advanced reinforcement learning post-training, we've achieved an optimal balance between semantic accuracy and visual excellence. The model demonstrates exceptional prompt adherence while delivering photorealistic imagery with stunning aesthetic quality and fine-grained details.
๐ญ
Intelligent World-Knowledge Reasoning:
The unified multimodal architecture endows HunyuanImage-3.0 with powerful reasoning capabilities. It leverages its extensive world knowledge to intelligently interpret user intent, automatically elaborating on sparse prompts with contextually appropriate details to produce superior, more complete visual outputs.
๐ ๏ธ Dependencies and Installation
๐ป System Requirements
๐ฅ๏ธ
Operating System:
Linux
๐ฎ
GPU:
NVIDIA GPU with CUDA support
๐พ
Disk Space:
170GB for model weights
๐ง
GPU Memory:
โฅ3ร80GB (4ร80GB recommended for better performance)
๐ฆ Environment Setup
๐
Python:
3.12+ (recommended and tested)
๐ฅ
PyTorch:
2.7.1
โก
CUDA:
12.8
๐ฅ Install Dependencies
# 1. First install PyTorch (CUDA 12.8 Version)
pip install torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu128
# 2. Then install other dependencies
pip install -r requirements.txt
Performance Optimizations
For
up to 3x faster inference
, install these optimizations:
# FlashAttention for faster attention computation
pip install flash-attn==2.8.3 --no-build-isolation
# FlashInfer for optimized moe inference. v0.3.1 is tested.
pip install flashinfer-python
๐ก
Installation Tips:
It is critical that the CUDA version used by PyTorch matches the system's CUDA version.
FlashInfer relies on this compatibility when compiling kernels at runtime. Pytorch 2.7.1+cu128 is tested.
GCC version >=9 is recommended for compiling FlashAttention and FlashInfer.
โก
Performance Tips:
These optimizations can significantly speed up your inference!
๐ก
Notation:
When FlashInfer is enabled, the first inference may be slower (about 10 minutes) due to kernel compilation. Subsequent inferences on the same machine will be much faster.
๐ Usage
๐ฅ Quick Start with Transformers
1๏ธโฃ Download model weights
# Download from HuggingFace and rename the directory.# Notice that the directory name should not contain dots, which may cause issues when loading using Transformers.
hf download tencent/HunyuanImage-3.0 --local-dir ./HunyuanImage-3
2๏ธโฃ Run with Transformers
from transformers import AutoModelForCausalLM
# Load the model
model_id = "./HunyuanImage-3"# Currently we can not load the model using HF model_id `tencent/HunyuanImage-3.0` directly # due to the dot in the name.
kwargs = dict(
attn_implementation="sdpa", # Use "flash_attention_2" if FlashAttention is installed
trust_remote_code=True,
torch_dtype="auto",
device_map="auto",
moe_impl="eager", # Use "flashinfer" if FlashInfer is installed
)
model = AutoModelForCausalLM.from_pretrained(model_id, **kwargs)
model.load_tokenizer(model_id)
# generate the image
prompt = "A brown and white dog is running on the grass"
image = model.generate_image(prompt=prompt, stream=True)
image.save("image.png")
๐ Local Installation & Usage
1๏ธโฃ Clone the Repository
git clone https://github.com/Tencent-Hunyuan/HunyuanImage-3.0.git
cd HunyuanImage-3.0/
2๏ธโฃ Download Model Weights
# Download from HuggingFace
hf download tencent/HunyuanImage-3.0 --local-dir ./HunyuanImage-3
3๏ธโฃ Run the Demo
python3 run_image_gen.py --model-id ./HunyuanImage-3 --verbose 1 --prompt "A brown and white dog is running on the grass"
4๏ธโฃ Command Line Arguments
Arguments
Description
Default
--prompt
Input prompt
(Required)
--model-id
Model path
(Required)
--attn-impl
Attention implementation. Either
sdpa
or
flash_attention_2
.
sdpa
--moe-impl
MoE implementation. Either
eager
or
flashinfer
eager
--seed
Random seed for image generation
None
--diff-infer-steps
Diffusion infer steps
50
--image-size
Image resolution. Can be
auto
, like
1280x768
or
16:9
auto
--save
Image save path.
image.png
--verbose
Verbose level. 0: No log; 1: log inference information.
0
๐จ Interactive Gradio Demo
Launch an interactive web interface for easy text-to-image generation.
1๏ธโฃ Install Gradio
pip install gradio>=4.21.0
2๏ธโฃ Configure Environment
# Set your model pathexport MODEL_ID="path/to/your/model"# Optional: Configure GPU usage (default: 0,1,2,3)export GPUS="0,1,2,3"# Optional: Configure host and port (default: 0.0.0.0:443)export HOST="0.0.0.0"export PORT="443"
3๏ธโฃ Launch the Web Interface
Basic Launch:
sh run_app.sh
With Performance Optimizations:
# Use both optimizations for maximum performance
sh run_app.sh --moe-impl flashinfer --attn-impl flash_attention_2
4๏ธโฃ Access the Interface
๐
Web Interface:
Open your browser and navigate to
http://localhost:443
(or your configured port)
Install performance extras (FlashAttention, FlashInfer) for faster inference.
MultiโGPU inference is recommended for the Base model.
๐ Prompt Guide
Manually Writing Prompts.
The Pretrain Checkpoint does not automatically rewrite or enhance input prompts, Instruct Checkpoint can rewrite or enhance input prompts with thinking . For optimal results currently, we recommend community partners consulting our official guide on how to write effective prompts.
We've included two system prompts in the PE folder of this repository that leverage DeepSeek to automatically enhance user inputs:
system_prompt_universal
: This system prompt converts photographic style, artistic prompts into a detailed one.
system_prompt_text_rendering
: This system prompt converts UI/Poster/Text Rending prompts to a deailed on that suits the model.
Note that these system prompts are in Chinese because Deepseek works better with Chinese system prompts. If you want to use it for English oriented model, you may translate it into English or refer to the comments in the PE file as a guide.
We also create a
Yuanqi workflow
to implent the universal one, you can directly try it.
Advanced Tips
Content Priority
: Focus on describing the main subject and action first, followed by details about the environment and style. A more general description framework is:
Main subject and scene + Image quality and style + Composition and perspective + Lighting and atmosphere + Technical parameters
. Keywords can be added both before and after this structure.
Image resolution
: Our model not only supports multiple resolutions but also offers both
automatic and specified resolution
options. In auto mode, the model automatically predicts the image resolution based on the input prompt. In specified mode (like traditional DiT), the model outputs an image resolution that strictly aligns with the user's chosen resolution.
More Cases
Our model can follow complex instructions to generate highโquality, creative images.
Our model can effectively process very long text inputs, enabling users to precisely control the finer details of generated images. Extended prompts allow for intricate elements to be accurately captured, making it ideal for complex projects requiring precision and creativity.
Show prompt
A cinematic medium shot captures a single Asian woman seated on a chair within a dimly lit room, creating an intimate and theatrical atmosphere. The composition is focused on the subject, rendered with rich colors and intricate textures that evoke a nostalgic and moody feeling.\n\nThe primary subject is a young Asian woman with a thoughtful and expressive countenance, her gaze directed slightly away from the camera. She is seated in a relaxed yet elegant posture on an ornate, vintage armchair. The chair is upholstered in a deep red velvet, its fabric showing detailed, intricate textures and slight signs of wear. She wears a simple, elegant dress in a dark teal hue, the material catching the light in a way that reveals its fine-woven texture. Her skin has a soft, matte quality, and the light delicately models the contours of her face and arms.\n\nThe surrounding room is characterized by its vintage decor, which contributes to the historic and evocative mood. In the immediate background, partially blurred due to a shallow depth of field consistent with a f/2.8 aperture, the wall is covered with wallpaper featuring a subtle, damask pattern. The overall color palette is a carefully balanced interplay of deep teal and rich red hues, creating a visually compelling and cohesive environment. The entire scene is detailed, from the fibers of the upholstery to the subtle patterns on the wall.\n\nThe lighting is highly dramatic and artistic, defined by high contrast and pronounced shadow play. A single key light source, positioned off-camera, projects gobo lighting patterns onto the scene, casting intricate shapes of light and shadow across the woman and the back wall. These dramatic shadows create a strong sense of depth and a theatrical quality. While some shadows are deep and defined, others remain soft, gently wrapping around the subject and preventing the loss of detail in darker areas. The soft focus on the background enhances the intimate feeling, drawing all attention to the expressive subject. The overall image presents a cinematic, photorealistic photography style.
Show prompt
A cinematic, photorealistic medium shot captures a high-contrast urban street corner, defined by the sharp intersection of light and shadow. The primary subject is the exterior corner of a building, rendered in a low-saturation, realistic style.\n\nThe building wall, which occupies the majority of the frame, is painted a warm orange with a finely detailed, rough stucco texture. Horizontal white stripes run across its surface. The base of the building is constructed from large, rough-hewn stone blocks, showing visible particles and texture. On the left, illuminated side of the building, there is a single window with closed, dark-colored shutters. Adjacent to the window, a simple black pendant lamp hangs from a thin, taut rope, casting a distinct, sharp-edged shadow onto the sunlit orange wall. The composition is split diagonally, with the right side of the building enveloped in a deep brown shadow. At the bottom of the frame, a smooth concrete sidewalk is visible, upon which the dynamic silhouette of a person is captured mid-stride, walking from right to left.\n\nIn the shallow background, the faint, out-of-focus outlines of another building and the bare, skeletal branches of trees are softly visible, contributing to the quiet urban atmosphere and adding a sense of depth to the scene. These elements are rendered with minimal detail to keep the focus on the foreground architecture.\n\nThe scene is illuminated by strong, natural sunlight originating from the upper left, creating a dramatic chiaroscuro effect. This hard light source casts deep, well-defined shadows, producing a sharp contrast between the brightly lit warm orange surfaces and the deep brown shadow areas. The lighting highlights the fine details in the wall texture and stone particles, emphasizing the photorealistic quality. The overall presentation reflects a high-quality photorealistic photography style, infused with a cinematic film noir aesthetic.
๐ค
SSAE (Machine Evaluation)
SSAE (Structured Semantic Alignment Evaluation) is an intelligent evaluation metric for image-text alignment based on advanced multimodal large language models (MLLMs). We extracted 3500 key points across 12 categories, then used multimodal large language models to automatically evaluate and score by comparing the generated images with these key points based on the visual content of the images. Mean Image Accuracy represents the image-wise average score across all key points, while Global Accuracy directly calculates the average score across all key points.
๐ฅ
GSB (Human Evaluation)
We adopted the GSB (Good/Same/Bad) evaluation method commonly used to assess the relative performance between two models from an overall image perception perspective. In total, we utilized 1,000 text prompts, generating an equal number of image samples for all compared models in a single run. For a fair comparison, we conducted inference only once for each prompt, avoiding any cherry-picking of results. When comparing with the baseline methods, we maintained the default settings for all selected models. The evaluation was performed by more than 100 professional evaluators.
๐ Citation
If you find HunyuanImage-3.0 useful in your research, please cite our work:
HunyuanImage-3.0-Instruct huggingface.co is an AI model on huggingface.co that provides HunyuanImage-3.0-Instruct's model effect (), which can be used instantly with this tencent HunyuanImage-3.0-Instruct model. huggingface.co supports a free trial of the HunyuanImage-3.0-Instruct model, and also provides paid use of the HunyuanImage-3.0-Instruct. Support call HunyuanImage-3.0-Instruct model through api, including Node.js, Python, http.
HunyuanImage-3.0-Instruct huggingface.co is an online trial and call api platform, which integrates HunyuanImage-3.0-Instruct's modeling effects, including api services, and provides a free online trial of HunyuanImage-3.0-Instruct, you can try HunyuanImage-3.0-Instruct online for free by clicking the link below.
tencent HunyuanImage-3.0-Instruct online free url in huggingface.co:
HunyuanImage-3.0-Instruct is an open source model from GitHub that offers a free installation service, and any user can find HunyuanImage-3.0-Instruct on GitHub to install. At the same time, huggingface.co provides the effect of HunyuanImage-3.0-Instruct install, users can directly use HunyuanImage-3.0-Instruct installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
HunyuanImage-3.0-Instruct install url in huggingface.co: