File size: 4,727 Bytes
cf77578 8290376 cf77578 c005c4b cf77578 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 | ---
license: apache-2.0
base_model: FireRedTeam/FireRedTTS3
pipeline_tag: text-to-speech
library_name: comfyui
tags:
- tts
- text-to-speech
- voice-cloning
- voice-design
- speech-editing
- fireredtts
- qwen3
- comfyui-custom-node
- int8
- convrot
- quantized
language:
- zh
- en
- yue
- ja
- ko
- es
- fr
- ru
- ar
- tr
- id
- pt
- it
- nl
- vi
- de
- uk
- th
- pl
- ro
- el
- cs
- fi
- hi
---
# FireRedTTS3-int8 (community INT8 ConvRot mirror)
INT8 ConvRot conversion of [FireRedTeam/FireRedTTS3](https://huggingface.co/FireRedTeam/FireRedTTS3) for
[FireRedTTS3-ComfyUI](https://github.com/Saganaki22/FireRedTTS3-ComfyUI), produced with the official
comfy-kitchen quantizer (`TensorWiseINT8Layout.quantize`, registry `quantize_int8_convrot_weight`).
Format per quantized Linear (current ComfyUI representation):
- `weight` — `torch.int8`, original `[out, in]` shape, contains the **offline Hadamard-rotated** weight (`W @ H^T` per 256-column group)
- `weight_scale` — `torch.float32`, `[out, 1]` per-output-row scale
- `bias` — original float bias
- `comfy_quant` — uint8 JSON: `{"format": "int8_tensorwise", "convrot": true, "convrot_groupsize": 256}`
At inference the companion custom node rotates activations online via
`comfy_kitchen.int8_linear(..., convrot=True, convrot_groupsize=256)` — dynamic per-row INT8 activation
quantization + INT8 GEMM, rescaled by `scale_x * scale_w`. No whole-weight dequantization on the hot path.
## What is quantized (safe profile, group size 256)
| Component | Quantized | Kept float |
| --- | --- | --- |
| `fireredtts3_base` | 321/332 Linears (1.73B params, 81.5% of core): all `backbone_llm.layers.*`, `patch_encoder.blocks.*`, `dit.blocks.*` | embeddings, norms, `spk_proj_*`, `patch_encoder.in_proj/out_proj`, `dit_head`, `dit.in_proj` (1600 % 256 != 0), `dit.t_embedder`, `dit.final_layer`, `stop_head`, Conv1d |
| `fireredtts3_instruct` | 321/331 Linears (1.73B params, 71.2% of core): same block families (`backbone_llm.model.layers.*`) | same exclusions |
| `redae` | nothing | everything |
| `campp` | nothing | everything |
## Sizes
| Core | Official fp32 | This repo |
| --- | --- | --- |
| `fireredtts3_base` | 8.48 GB | 3.30 GB |
| `fireredtts3_instruct` | 8.48 GB | 3.30 GB |
| `redae` / `campp` / tokenizer | copied through unchanged | |
## Validation (base variant, full suite; instruct smoke-tested)
- Per-layer weight roundtrip (official quantize -> official dequantize): worst rel-L2 **0.00967**, worst cosine **0.999953** over 321 layers
- Real-activation comparison vs fp32 through the same `comfy_kitchen.int8_linear` runtime: worst rel-L2 **0.01162**, worst cosine **0.999932**
- On-disk structure: all 677 original keys preserved, scales fp32 `[N,1]` and positive, Conv1d/RedAE/CAM++ untouched
- Runtime proof: 321 `ConvRotInt8Linear` modules, >42k counted INT8 ConvRot kernel calls during generation, weights stay int8 across unload/reload
- BF16 vs INT8 generation (same seed/settings): identical patch counts (200/200), finite latents, EN/ZH ASR-verified, speaker-similarity parity (0.9007 vs 0.8991)
- Peak VRAM 13.1 -> 8.3 GiB; generation ~1.3x slower (memory optimization, honestly reported)
## Usage Disclaimer
- The project incorporates zero-shot voice cloning functionality; Please note that this capability is intended **solely for academic research purposes**.
- **DO NOT** use this model for **ANY illegal activities**❗️❗️
- The developers assume no liability for any misuse of this model.
- If you identify any instances of **abuse**, **misuse**, or **fraudulent** activities related to this project, **please report them to our team immediately.**
## Citation
```bib
@article{fireredtts3,
title = {FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations},
author = {FireRed Team},
journal = {arXiv preprint},
year = {2026},
}
```
## Acknowledgements
- [Qwen3](https://github.com/QwenLM/Qwen3) and [Qwen2-Audio](https://github.com/QwenLM/Qwen2-Audio) for the language model and audio understanding foundations
- [DiTAR](https://arxiv.org/abs/2502.03930) for the patch-level diffusion autoregressive formulation
- [X-Codec](https://github.com/zhenye234/xcodec) for the discriminator design used in RedAE training
- [CAM++](https://modelscope.cn/models/iic/speech_campplus_sv_en_voxceleb_16k) for speaker embedding extraction
- [fastText](https://fasttext.cc/docs/en/language-identification.html) for automatic language identification
- [WeTextProcessing](https://github.com/wenet-e2e/WeTextProcessing) (wetext) for the Chinese / English text normalization front-end
All credit to the FireRed Team — see the upstream repo and model card. Apache-2.0.
|