GGUF Quantization support for native ComfyUI models
This is currently very much WIP. These custom nodes provide support for model files stored in the GGUF format popularized by
llama.cpp
.
While quantization wasn't feasible for regular UNET models (conv2d), transformer/DiT models such as flux seem less affected by quantization. This allows running it in much lower bits per weight variable bitrate quants on low-end GPUs. For further VRAM savings, a node to load a quantized version of the T5 text encoder is also included.
Note: The "Force/Set CLIP Device" is
NOT
part of this node pack. Do not install it if you only have one GPU. Do not set it to cuda:0 then complain about OOM errors if you do not undestand what it is for. There is not need to copy the workflow above, just use your own workflow and replace the stock "Load Diffusion Model" with the "Unet Loader (GGUF)" node.
Installation
Make sure your ComfyUI is on a recent-enough version to support custom ops when loading the UNET-only.
To install the custom node normally, git clone this repository into your custom nodes folder (
ComfyUI/custom_nodes
) and install the only dependency for inference (
pip install --upgrade gguf
)
git clone https://github.com/city96/ComfyUI-GGUF
To install the custom node on a standalone ComfyUI release, open a CMD inside the "ComfyUI_windows_portable" folder (where your
run_nvidia_gpu.bat
file is) and use the following commands:
On MacOS sequoia, torch 2.4.1 seems to be required, as 2.6.X nightly versions cause a "M1 buffer is not large enough" error. See
this issue
for more information/workarounds.
Usage
Simply use the GGUF Unet loader found under the
bootleg
category. Place the .gguf model files in your
ComfyUI/models/unet
folder.
LoRA loading is experimental but it should work with just the built-in LoRA loader node(s).
Initial support for quantizing T5 has also been added recently, these can be used using the various
*CLIPLoader (gguf)
nodes which can be used inplace of the regular ones. For the CLIP model, use whatever model you were using before for CLIP. The loader can handle both types of files -
gguf
and regular
safetensors
/
bin
.
See the instructions in the
tools
folder for how to create your own quants.
Runs of Aero-Ex ComfyUI-GGUF on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About ComfyUI-GGUF huggingface.co Model
ComfyUI-GGUF huggingface.co
ComfyUI-GGUF huggingface.co is an AI model on huggingface.co that provides ComfyUI-GGUF's model effect (), which can be used instantly with this Aero-Ex ComfyUI-GGUF model. huggingface.co supports a free trial of the ComfyUI-GGUF model, and also provides paid use of the ComfyUI-GGUF. Support call ComfyUI-GGUF model through api, including Node.js, Python, http.
ComfyUI-GGUF huggingface.co is an online trial and call api platform, which integrates ComfyUI-GGUF's modeling effects, including api services, and provides a free online trial of ComfyUI-GGUF, you can try ComfyUI-GGUF online for free by clicking the link below.
Aero-Ex ComfyUI-GGUF online free url in huggingface.co:
ComfyUI-GGUF is an open source model from GitHub that offers a free installation service, and any user can find ComfyUI-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of ComfyUI-GGUF install, users can directly use ComfyUI-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.