INT8 ConvRot conversion of
FireRedTeam/FireRedTTS3
for
FireRedTTS3-ComfyUI
, produced with the official
comfy-kitchen quantizer (
TensorWiseINT8Layout.quantize
, registry
quantize_int8_convrot_weight
).
Format per quantized Linear (current ComfyUI representation):
weight
—
torch.int8
, original
[out, in]
shape, contains the
offline Hadamard-rotated
weight (
W @ H^T
per 256-column group)
At inference the companion custom node rotates activations online via
comfy_kitchen.int8_linear(..., convrot=True, convrot_groupsize=256)
— dynamic per-row INT8 activation
quantization + INT8 GEMM, rescaled by
scale_x * scale_w
. No whole-weight dequantization on the hot path.
What is quantized (safe profile, group size 256)
Component
Quantized
Kept float
fireredtts3_base
321/332 Linears (1.73B params, 81.5% of core): all
backbone_llm.layers.*
,
patch_encoder.blocks.*
,
dit.blocks.*
FireRedTTS3-int8 huggingface.co is an AI model on huggingface.co that provides FireRedTTS3-int8's model effect (), which can be used instantly with this drbaph FireRedTTS3-int8 model. huggingface.co supports a free trial of the FireRedTTS3-int8 model, and also provides paid use of the FireRedTTS3-int8. Support call FireRedTTS3-int8 model through api, including Node.js, Python, http.
FireRedTTS3-int8 huggingface.co is an online trial and call api platform, which integrates FireRedTTS3-int8's modeling effects, including api services, and provides a free online trial of FireRedTTS3-int8, you can try FireRedTTS3-int8 online for free by clicking the link below.
drbaph FireRedTTS3-int8 online free url in huggingface.co:
FireRedTTS3-int8 is an open source model from GitHub that offers a free installation service, and any user can find FireRedTTS3-int8 on GitHub to install. At the same time, huggingface.co provides the effect of FireRedTTS3-int8 install, users can directly use FireRedTTS3-int8 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.