File size: 4,727 Bytes
cf77578
 
 
8290376
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cf77578
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c005c4b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cf77578
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
---
license: apache-2.0
base_model: FireRedTeam/FireRedTTS3
pipeline_tag: text-to-speech
library_name: comfyui
tags:
- tts
- text-to-speech
- voice-cloning
- voice-design
- speech-editing
- fireredtts
- qwen3
- comfyui-custom-node
- int8
- convrot
- quantized
language:
- zh
- en
- yue
- ja
- ko
- es
- fr
- ru
- ar
- tr
- id
- pt
- it
- nl
- vi
- de
- uk
- th
- pl
- ro
- el
- cs
- fi
- hi
---

# FireRedTTS3-int8 (community INT8 ConvRot mirror)

INT8 ConvRot conversion of [FireRedTeam/FireRedTTS3](https://huggingface.co/FireRedTeam/FireRedTTS3) for
[FireRedTTS3-ComfyUI](https://github.com/Saganaki22/FireRedTTS3-ComfyUI), produced with the official
comfy-kitchen quantizer (`TensorWiseINT8Layout.quantize`, registry `quantize_int8_convrot_weight`).

Format per quantized Linear (current ComfyUI representation):

- `weight``torch.int8`, original `[out, in]` shape, contains the **offline Hadamard-rotated** weight (`W @ H^T` per 256-column group)
- `weight_scale``torch.float32`, `[out, 1]` per-output-row scale
- `bias` — original float bias
- `comfy_quant` — uint8 JSON: `{"format": "int8_tensorwise", "convrot": true, "convrot_groupsize": 256}`

At inference the companion custom node rotates activations online via
`comfy_kitchen.int8_linear(..., convrot=True, convrot_groupsize=256)` — dynamic per-row INT8 activation
quantization + INT8 GEMM, rescaled by `scale_x * scale_w`. No whole-weight dequantization on the hot path.

## What is quantized (safe profile, group size 256)

| Component | Quantized | Kept float |
| --- | --- | --- |
| `fireredtts3_base` | 321/332 Linears (1.73B params, 81.5% of core): all `backbone_llm.layers.*`, `patch_encoder.blocks.*`, `dit.blocks.*` | embeddings, norms, `spk_proj_*`, `patch_encoder.in_proj/out_proj`, `dit_head`, `dit.in_proj` (1600 % 256 != 0), `dit.t_embedder`, `dit.final_layer`, `stop_head`, Conv1d |
| `fireredtts3_instruct` | 321/331 Linears (1.73B params, 71.2% of core): same block families (`backbone_llm.model.layers.*`) | same exclusions |
| `redae` | nothing | everything |
| `campp` | nothing | everything |

## Sizes

| Core | Official fp32 | This repo |
| --- | --- | --- |
| `fireredtts3_base` | 8.48 GB | 3.30 GB |
| `fireredtts3_instruct` | 8.48 GB | 3.30 GB |
| `redae` / `campp` / tokenizer | copied through unchanged | |

## Validation (base variant, full suite; instruct smoke-tested)

- Per-layer weight roundtrip (official quantize -> official dequantize): worst rel-L2 **0.00967**, worst cosine **0.999953** over 321 layers
- Real-activation comparison vs fp32 through the same `comfy_kitchen.int8_linear` runtime: worst rel-L2 **0.01162**, worst cosine **0.999932**
- On-disk structure: all 677 original keys preserved, scales fp32 `[N,1]` and positive, Conv1d/RedAE/CAM++ untouched
- Runtime proof: 321 `ConvRotInt8Linear` modules, >42k counted INT8 ConvRot kernel calls during generation, weights stay int8 across unload/reload
- BF16 vs INT8 generation (same seed/settings): identical patch counts (200/200), finite latents, EN/ZH ASR-verified, speaker-similarity parity (0.9007 vs 0.8991)
- Peak VRAM 13.1 -> 8.3 GiB; generation ~1.3x slower (memory optimization, honestly reported)

## Usage Disclaimer

- The project incorporates zero-shot voice cloning functionality; Please note that this capability is intended **solely for academic research purposes**.
- **DO NOT** use this model for **ANY illegal activities**❗️❗️
- The developers assume no liability for any misuse of this model.
- If you identify any instances of **abuse**, **misuse**, or **fraudulent** activities related to this project, **please report them to our team immediately.**


## Citation

```bib
@article{fireredtts3,
  title   = {FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations},
  author  = {FireRed Team},
  journal = {arXiv preprint},
  year    = {2026},
}
```


## Acknowledgements

- [Qwen3](https://github.com/QwenLM/Qwen3) and [Qwen2-Audio](https://github.com/QwenLM/Qwen2-Audio) for the language model and audio understanding foundations
- [DiTAR](https://arxiv.org/abs/2502.03930) for the patch-level diffusion autoregressive formulation
- [X-Codec](https://github.com/zhenye234/xcodec) for the discriminator design used in RedAE training
- [CAM++](https://modelscope.cn/models/iic/speech_campplus_sv_en_voxceleb_16k) for speaker embedding extraction
- [fastText](https://fasttext.cc/docs/en/language-identification.html) for automatic language identification
- [WeTextProcessing](https://github.com/wenet-e2e/WeTextProcessing) (wetext) for the Chinese / English text normalization front-end


All credit to the FireRed Team — see the upstream repo and model card. Apache-2.0.