Figure 1: We address the reconstruction accuracy drop of high spatial-compression autoencoders.
Figure 2: DC-AE delivers significant training and inference speedup without performance drop.
Figure 3: DC-AE enables efficient text-to-image generation on the laptop.
Abstract
We present Deep Compression Autoencoder (DC-AE), a new family of autoencoder models for accelerating high-resolution diffusion models. Existing autoencoder models have demonstrated impressive results at a moderate spatial compression ratio (e.g., 8x), but fail to maintain satisfactory reconstruction accuracy for high spatial compression ratios (e.g., 64x). We address this challenge by introducing two key techniques: (1) Residual Autoencoding, where we design our models to learn residuals based on the space-to-channel transformed features to alleviate the optimization difficulty of high spatial-compression autoencoders; (2) Decoupled High-Resolution Adaptation, an efficient decoupled three-phases training strategy for mitigating the generalization penalty of high spatial-compression autoencoders. With these designs, we improve the autoencoder's spatial compression ratio up to 128 while maintaining the reconstruction quality. Applying our DC-AE to latent diffusion models, we achieve significant speedup without accuracy drop. For example, on ImageNet 512x512, our DC-AE provides 19.1x inference speedup and 17.9x training speedup on H100 GPU for UViT-H while achieving a better FID, compared with the widely used SD-VAE-f8 autoencoder.
If DC-AE is useful or relevant to your research, please kindly recognize our contributions by citing our papers:
@article{chen2024deep,
title={Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models},
author={Chen, Junyu and Cai, Han and Chen, Junsong and Xie, Enze and Yang, Shang and Tang, Haotian and Li, Muyang and Lu, Yao and Han, Song},
journal={arXiv preprint arXiv:2410.10733},
year={2024}
}
Runs of mit-han-lab dc-ae-lite-f32c32-sana-1.1-diffusers on huggingface.co
108
Total runs
6
24-hour runs
15
3-day runs
34
7-day runs
74
30-day runs
More Information About dc-ae-lite-f32c32-sana-1.1-diffusers huggingface.co Model
dc-ae-lite-f32c32-sana-1.1-diffusers huggingface.co is an AI model on huggingface.co that provides dc-ae-lite-f32c32-sana-1.1-diffusers's model effect (), which can be used instantly with this mit-han-lab dc-ae-lite-f32c32-sana-1.1-diffusers model. huggingface.co supports a free trial of the dc-ae-lite-f32c32-sana-1.1-diffusers model, and also provides paid use of the dc-ae-lite-f32c32-sana-1.1-diffusers. Support call dc-ae-lite-f32c32-sana-1.1-diffusers model through api, including Node.js, Python, http.
dc-ae-lite-f32c32-sana-1.1-diffusers huggingface.co is an online trial and call api platform, which integrates dc-ae-lite-f32c32-sana-1.1-diffusers's modeling effects, including api services, and provides a free online trial of dc-ae-lite-f32c32-sana-1.1-diffusers, you can try dc-ae-lite-f32c32-sana-1.1-diffusers online for free by clicking the link below.
mit-han-lab dc-ae-lite-f32c32-sana-1.1-diffusers online free url in huggingface.co:
dc-ae-lite-f32c32-sana-1.1-diffusers is an open source model from GitHub that offers a free installation service, and any user can find dc-ae-lite-f32c32-sana-1.1-diffusers on GitHub to install. At the same time, huggingface.co provides the effect of dc-ae-lite-f32c32-sana-1.1-diffusers install, users can directly use dc-ae-lite-f32c32-sana-1.1-diffusers installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
dc-ae-lite-f32c32-sana-1.1-diffusers install url in huggingface.co: