Readme
NitroFusion: High-Fidelity Single-Step Diffusion through Dynamic Adversarial Training
License
NitroSD-Realism is released under cc-by-nc-4.0 , following its base model DMD2 .
NitroSD-Vibrant is released under openrail++ .

High-Fidelity Single-Step Diffusion through Dynamic Adversarial Training
Fast sdxl with higher quality
CogVLM2: Visual Language Models for Image and Video Understanding
Depth estimation with faster inference speed, fewer parameters, and higher depth accuracy.
Updated to OpenVoice v2: Versatile Instant Voice Cloning
Audio-based Lip Synchronization for Talking Head Video
Fast and High-Quality Text-to-video Generation
OmniGen: Unified Image Generation
Scalable Streaming Speech Synthesis with Large Language Models
DiT-based video generation model for generating high-quality videos in real-time
Convert LLM's coding to image generation
Sharp Monocular Metric Depth in Less Than a Second
Minimal and Universal Control for Diffusion Transformer - demo for Subject-driven generation
Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
CogVLM2: Visual Language Models for Image and Video Understanding
Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Extended video synthesis model that generates 128 frames
Depth Any Video with Scalable Synthetic Data
Generating Consistent Long Depth Sequences for Open-world Videos
One Diffusion to Generate Them All
Diffusion-based Visual Foundation Model for High-quality Dense Prediction
Efficient Visual Generation with Hybrid Autoregressive Transformer
Minimal and Universal Control for Diffusion Transformer - demo for Spatially aligned control
Image-to-Video Diffusion Models with An Expert Transformer
Finer and Faster Text-to-Image Generation via Relay Diffusion
Text-to-Video Diffusion Models with An Expert Transformer
Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
Autoregressive Video Generation without Vector Quantization
Emu3-Gen for image generation
Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution
Emu3-Chat for vision-language understanding
Autoregressive Image Generation without Vector Quantization
Let Vision Language Models Reason Step-by-Step
Enhancing Video Model Post-Training through Data, Reward, and Conditional Guidance Design
Text-to-Video Diffusion Models with An Expert Transformer