This is a BF16 conversion of LongCat-AudioDiT-3.5B — a state-of-the-art diffusion-based zero-shot TTS model by Meituan that operates directly in the waveform latent space. Converting from FP32 to BF16 halves the on-disk size and VRAM usage with negligible quality loss, making it the recommended variant for most users.
Original (3.5B FP32)
This (3.5B BF16)
Weight dtype
float32
bfloat16
Activation dtype
float32
bfloat16
File size
~14 GB
~7 GB
VRAM (inference)
~20 GB
~12 GB
Quality
Reference
Virtually identical
Extra dependencies
none
none
Conversion Details
All model weights — DiT transformer backbone, Wav-VAE, and text encoder — are converted from float32 to bfloat16. BF16 preserves the same dynamic range as FP32 (8 exponent bits) while halving memory usage, making it the lossless practical choice for inference on modern GPUs.
No post-training quantization, calibration data, or scale factors are required. The model is a direct dtype cast and is fully compatible with the original
audiodit
inference code.
Hardware Requirements
GPU:
NVIDIA GPU with CUDA support (BF16 supported on Ampere and newer; falls back gracefully on older hardware)
VRAM:
~7 GB
CPU:
Supported but slow
Usage — ComfyUI (Recommended)
The easiest way to use this model is with
ComfyUI-LongCat-AudioDIT-TTS
, which has native support for this BF16 model with zero extra setup.
Installation
Install the ComfyUI node via
ComfyUI Manager
(search
LongCat-AudioDiT
) or manually:
cd ComfyUI/custom_nodes
git clone https://github.com/Saganaki22/ComfyUI-LongCat-AudioDIT-TTS.git
The model
auto-downloads on first use
— select
LongCat-AudioDiT-3.5B-bf16
from the model dropdown in any LongCat node.
dtype
:
auto
or
bf16
— matches this model's native dtype
guidance_method
:
cfg
for TTS,
apg
for voice cloning
steps
:
16
(balanced),
32
(higher quality)
keep_model_loaded
:
True
for repeated use
This is the recommended variant
for most users — best balance of quality, VRAM usage, and compatibility.
About LongCat-AudioDiT
LongCat-AudioDiT is a non-autoregressive diffusion-based TTS model from Meituan that achieves state-of-the-art zero-shot voice cloning performance on the Seed benchmark. Unlike previous methods relying on mel-spectrograms, it operates directly in the waveform latent space using only a Wav-VAE and a DiT backbone.
The 3.5B variant achieves
0.818 SIM on Seed-ZH
and
0.797 SIM on Seed-Hard
, surpassing both open-source and closed-source competitors.
LongCat-AudioDiT-3.5B-bf16 huggingface.co is an AI model on huggingface.co that provides LongCat-AudioDiT-3.5B-bf16's model effect (), which can be used instantly with this drbaph LongCat-AudioDiT-3.5B-bf16 model. huggingface.co supports a free trial of the LongCat-AudioDiT-3.5B-bf16 model, and also provides paid use of the LongCat-AudioDiT-3.5B-bf16. Support call LongCat-AudioDiT-3.5B-bf16 model through api, including Node.js, Python, http.
LongCat-AudioDiT-3.5B-bf16 huggingface.co is an online trial and call api platform, which integrates LongCat-AudioDiT-3.5B-bf16's modeling effects, including api services, and provides a free online trial of LongCat-AudioDiT-3.5B-bf16, you can try LongCat-AudioDiT-3.5B-bf16 online for free by clicking the link below.
drbaph LongCat-AudioDiT-3.5B-bf16 online free url in huggingface.co:
LongCat-AudioDiT-3.5B-bf16 is an open source model from GitHub that offers a free installation service, and any user can find LongCat-AudioDiT-3.5B-bf16 on GitHub to install. At the same time, huggingface.co provides the effect of LongCat-AudioDiT-3.5B-bf16 install, users can directly use LongCat-AudioDiT-3.5B-bf16 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
LongCat-AudioDiT-3.5B-bf16 install url in huggingface.co: