File size: 19,237 Bytes
405c697 b22270e 405c697 b22270e 405c697 b22270e 405c697 38f5ab7 c9ecb07 38f5ab7 b22270e c9ecb07 b22270e 405c697 b22270e 405c697 09e8a57 fa92dce cf72522 38f5ab7 cf72522 301e170 930003a 3ad028e 3506ed7 38f5ab7 3506ed7 38f5ab7 3506ed7 38f5ab7 3506ed7 fa92dce 3506ed7 09e8a57 9636ed0 405c697 9636ed0 405c697 b5560c8 38f5ab7 b5560c8 405c697 f05c126 b5560c8 38f5ab7 b5560c8 405c697 f05c126 405c697 b5560c8 3d36930 b5560c8 405c697 ac80cba 38f5ab7 ac80cba f05c126 405c697 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 | ---
license: other
license_name: ltx-2.5-community
tags:
- quantization
- 4-bit
- int4
- nvfp4
- gptq
- awq
- text-encoder
- text-to-video
- image-to-video
- video-to-video
- ltx-video
- low-vram
- comfyui
base_model: Lightricks/LTX-2.5
---
[](https://buymeacoffee.com/choijjs83q)
# LTX-2.5 Text Encoder β 4-bit, 8 GB
> **The 8 GB in the name is the file**, not the card. On disk this is 8.46 GB
> against the original's 26.264 GB. Running it needs about **9.7 GB of VRAM** β
> an 8 GB card is not enough. See [Memory](#memory--what-the-card-actually-needs).
> **EN** β I'm a student researching ML quantization. Getting this one done
> burned through so much in server bills that from here on I'll only be able to
> afford *niu lai* movies. Thank you for using the model.
>
> **νκ΅μ΄** β μ λ ML μμνλ₯Ό μ°κ΅¬νλ νμμ
λλ€. μμνλ₯Ό μ§ννλ©΄μ μλ²
> λΉμ©μ λ무 λ§μ΄ μ¨μ, μμΌλ‘ μνλ *niu lai*λ§ λ΄μΌ ν κ² κ°μ΅λλ€. λͺ¨λΈμ
> μ¬μ©ν΄ μ£Όμ
μ κ°μ¬ν©λλ€.
>
> **δΈζ** β ζζ―δΈεη η©ΆζΊε¨ε¦δΉ ιεηε¦ηγεθΏζ¬‘ιεη§ζδΊε€ͺε€ζε‘ε¨θ΄Ήη¨οΌ
> δ»₯εηη΅ε½±ε€§ζ¦εͺθ½η *niu lai* δΊγζθ°’ζ¨δ½Ώη¨θΏδΈͺ樑εγ
>
> β **[Buy me a coffee](https://buymeacoffee.com/choijjs83q)** Β· [μ»€νΌ ν μ μ¬μ£ΌκΈ°](https://buymeacoffee.com/choijjs83q) Β· [θ―·ζεζ―εε‘](https://buymeacoffee.com/choijjs83q)
The Gemma4-12B text encoder that LTX-2.5 needs in order to read a prompt,
compressed from **26.264 GB to 8.46 GB** (3.10x) and runnable on **any CUDA
GPU** β no minimum compute capability, no custom kernels, no CUDA 13.
If you have been unable to run LTX-2.5 because the text encoder alone wanted
26 GB, this is the part that was in your way. Drop it in and the rest of the
model is unchanged.
Built from the encoder Lightricks published on 2026-08-17
(`1b92891c`, *"Aligns the published encoders with the LTX-2.5 model
checkpoints"*) β the revision that matches the released DiT.
## Why this one
Every other public quantization of this encoder needs recent hardware:
| build | size | needs |
|---|---:|---|
| Lightricks BF16 | 26.264 GB | β |
| Lightricks `comfy-int8-convrot` | 15.373 GB | cc 8.9 |
| DmitryDB `nvfp4` | 11.197 GB | cc 8.9, `comfy_kitchen`, CUDA 13 |
| joeygambino / Winnougan `w4a8` | 10.604 GB | SM 8.0+ |
| **this** | **8.46 GB** | **nothing beyond PyTorch** |
Dequantization happens on the CPU at load and the resident model is BF16, so
there is no kernel requirement to satisfy. Verified running on **Tesla V100
(cc 7.0)**, **A100 (cc 8.0)** and **RTX 6000 Ada (cc 8.9)**.
It is also the smallest of the set, because it quantizes the three tensors the
other nvfp4 recipe protects as BF16 "precision islands" β `embed_tokens` and
both aggregate tables, 4.4 GB of the source.
## Files
| file | build | video relL2 | audio relL2 | `\|\|Q\|\|/\|\|W\|\|` |
|---|---|---:|---:|---:|
| `A3.packed.safetensors` | bypass guard + group-bounded AWQ | **0.05204** | **0.04827** | **4.216** |
| `A0.packed.safetensors` | legacy guard, for comparison | 0.06061 | 0.06095 | 302.654 |
**Use `A3`.** `A0` is published only so the comparison can be checked; it
carries weights up to 300x their proper norm in near-dead channels, which is
harmless on this calibration set and fragile by construction.
## Samples
`samples/` holds every clip twice β once from the **BF16 original (26.264 GB)**
and once from this **4-bit build**, with everything downstream of the encoder
held identical: same DiT, same seed, same schedule, same VAE settings, one
process. `compare-NN.mp4` stacks each pair, BF16 on the left.
Two sets, because they answer different questions. `compare-NN.mp4` uses the
vendor's `euler_ancestral`, which is what you actually get; it re-rolls noise
every step, so the two builds return different *takes*. `compare-det-NN.mp4`
uses deterministic `euler`, where the seed fixes the starting noise and the
conditioning is the only thing left that can move a pixel β that is the set
that attributes a difference to the encoder.
Mean absolute error against BF16 over every decoded frame:
| clip | `euler_ancestral` | `euler` |
|---|---:|---:|
| robot in rain | 0.0172 | 0.0182 |
| dune | 0.0596 | 0.0580 |
| forge | 0.0098 | 0.0129 |
| night road | 0.0702 | 0.0252 |
| smoke | 0.0128 | 0.0165 |
**Four of the five deterministic pairs hold together** β same composition, same
lighting, same timing, differing in surface detail. **The dune does not**: BF16
renders a soldier in fatigues where the 4-bit renders a man in a business suit,
from the same seed under a deterministic sampler. That difference belongs to
the encoder, and it is published rather than cropped out. The qualification it
deserves is that neither build followed that prompt β it asked for an astronaut
and got neither β so the model had no confident answer there for a small
conditioning change to disturb.
All five deterministic pairs at a glance β BF16 left, 4-bit right, one row per
prompt. Rows 1, 3, 4 and 5 hold together; row 2, the dune, is where the
compression is visible.

### 15 seconds, 1024x640, generated in the Space
<video controls width="100%" poster="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/moe-idol-poster.png" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/moe-idol-15s.mp4"></video>

361 frames with sound, from a 300-word prompt naming ten sequenced gestures, made
entirely through [the Space](https://huggingface.co/spaces/topabaem/LTX-2.5-Text-Encoder-4bit-8GB-Demo)
on ZeroGPU in 198 s. The full prompt is in
[`samples/idol/prompt.txt`](https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/idol/prompt.txt).
Two things had to be fixed before a clip this long held together, and both were
in this project's own defaults rather than in the compression:
* **The sampling schedule did not adapt to length.** `DISTILLED_SIGMAS` is nine
fixed numbers; `LTXVScheduler` derives its shift from the latent's token count,
and a long clip carries several times more. Left fixed, the sample never
converges β furniture renders semi-transparent and saturation halves. Clips
past 15 s now switch automatically.
* **The vendor's second pass was being skipped**, on an unmeasured assumption
that a 16 GB card could not afford it. It can, at 10.03 GiB, and it is worth
4.1x the Laplacian variance.
Details, the measurements, and the failures behind both are in
[`samples/long/schedule.md`](https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/long/schedule.md).
### The samples, as the vendor's pipeline actually renders them

Six clips at **1024x640** β sample at 512x320, upscale the latent 2x, sample
again. That second pass had been skipped here from the start on an assumption
that a 16 GB card could not afford it; measured, it peaks at 10.03 GiB, and it
is worth **4.1x the Laplacian variance**. Everything else on this page is
one-pass and softer for it. [`samples/sharp/`](https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/sharp/README.md), including
the one of the six that does not follow its prompt and why.
### Fifteen seconds, on one 16 GB card
<video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/idol/idol-15s.mp4"></video>
353 frames at 512x320, generated in 123 s at a 6.26 GiB peak β no more memory
than the ten-second version, so the length ceiling was not reached. The prompt
and an honest account of what it did *not* do β the audio is not speech, and the
"2D anime" instruction is ignored at other lengths by the BF16 original too β
are in [`samples/idol/`](https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/idol/README.md).
### Watch the two that matter
**Forge β the pair holding together.** Deterministic sampler, same seed. Left is
BF16, right is 4-bit.
<video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/compare-det-02.mp4"></video>
**Dune β the pair that does not.** Same terrain, same sun, same shadow, same
walk, and a different person.
<video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/compare-det-01.mp4"></video>
Individually, the forge under the vendor's own sampler:
| BF16 original | 4-bit |
|---|---|
| <video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/bf16-02.mp4"></video> | <video controls width="100%" src="https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/samples/4bit-02.mp4"></video> |
`samples/README.md` carries the prompts, the per-clip conditioning drift, and
what these clips do and do not establish. All thirty clips are in `samples/`,
and the Space plays them side by side under its **BF16 vs 4-bit** tab.
## Memory β what the card actually needs
| | resident (default) | dequantized |
|---|---:|---:|
| PyTorch allocated, peak | 8.48 GiB | ~22.3 GiB |
| PyTorch reserved, peak | 9.33 GiB | β |
| **what `nvidia-smi` shows** | **9.70 GiB** | β |
**Size your card from the last row.** The first is `max_memory_allocated`, which
counts only live allocator blocks β it misses the CUDA context and everything
the caching allocator reserved and has not handed back, and it undercounts by
nearly 2 GiB here. Earlier versions of this card quoted that number, and anyone
who bought an 8 GB card on the strength of it would have been wrong.
Measured on a **Tesla V100-SXM2-16GB (cc 7.0)** with a 55-token prompt, through
the ComfyUI-faithful path that left-pads to 1024 tokens. So:
* **16 GB and up** β comfortable.
* **12 GB** β fits.
* **10 GB** β fits, with roughly 300 MB of headroom. Nothing else on the card.
* **8 GB** β does not fit.
The two modes produce `torch.equal` conditioning, so the choice is footprint
against speed and never quality. Resident dequantizes inside `forward`, which
costs time on short prompts; dequantized builds one dense BF16 model at load and
then needs a card that can hold 26 GB.
These figures are for the **text encoder alone**. Generating video also needs the
DiT and the VAEs, which this repository does not contain β the pipeline in the
Space loads a `Q3_K_M` DiT alongside it.
## Use it
### Standalone
Five packages and one file. No build step, no custom CUDA kernels, no
compilation.
```bash
pip install -r <(curl -sL https://huggingface.co/topabaem/LTX-2.5-Text-Encoder-4bit-8GB/resolve/main/requirements.txt)
```
```python
from huggingface_hub import hf_hub_download, snapshot_download
repo = "topabaem/LTX-2.5-Text-Encoder-4bit-8GB"
# The loader ships with the weights; put it on the path before importing it.
import sys, os
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "ltx_packed_codec.py")))
from ltx_packed_codec import load_packed_model
from transformers import AutoTokenizer
packed = hf_hub_download(repo, "A3.packed.safetensors")
encoder_dir = snapshot_download(repo, allow_patterns=["encoder-hf/*"]) + "/encoder-hf"
model = load_packed_model(encoder_dir, packed, resident=True)
tokenizer = AutoTokenizer.from_pretrained(encoder_dir)
```
Verified end to end in a clean virtualenv containing nothing but those five
packages, on **torch 2.13.0 / transformers 5.15.1** and on **torch 2.10.0 /
transformers 5.12.1**. `encoder-hf/` is config and tokenizer only, 31 MB β the
26 GB original is not needed.
### What your GPU has to support
Nothing unusual. The format needs no fp8 hardware: the group scales are stored
as `float8_e4m3fn` bytes and converted in software during a CPU-side decode, so
`float8` here is a container and never an instruction. There is no minimum
compute capability, no `comfy_kitchen`, no CUDA 13. The resident model is BF16
and runs on cards with no bf16 tensor cores at all.
**The one real constraint is your torch wheel, not your GPU.** The default wheel
on PyPI is now a cu130 build and cu130 dropped Volta. Measured on a V100:
| torch build | device | result |
|---|---|---|
| 2.13.0+cu130 | CPU | works |
| 2.13.0+cu130 | V100, sm_70 | **`no kernel image is available`** |
| 2.10.0+cu128 | V100, sm_70 | works |
That failure arrives at the *first kernel launch*, well after
`torch.cuda.is_available()` has returned `True`, so it does not look like an
installation problem β which is why **the loader checks first**. Before reading
a byte of the 8.46 GB it compares your card against
`torch.cuda.get_arch_list()` and, on a mismatch, stops with the fix:
```
this torch (2.13.0+cu130) has no kernels for Tesla V100-SXM2-16GB (sm_70).
It was built for sm_75, sm_80, sm_86, sm_90, sm_100, sm_120, and the first CUDA
op would fail with 'no kernel image is available for execution on the device'.
The model is fine - it needs no custom kernels. Install a torch built for your
card, e.g. for sm_70:
pip install torch --index-url https://download.pytorch.org/whl/cu128
or pass device='cpu' to load without touching the GPU.
```
On sm_70 install a cu128 build:
```bash
pip install torch --index-url https://download.pytorch.org/whl/cu128
```
Ampere and newer are unaffected β the stock wheel carries kernels for them.
### ComfyUI
`comfy_nodes/ComfyUI-LTXPacked/` provides **LTX Packed Encoder Loader** and
**LTX Packed Text Encode**, which replace `CLIPTextEncode` and emit a
CONDITIONING directly.
They do not go through `LTXAVTextEncoderLoader`, and cannot: ComfyUI's LTX CLIP
path wants a sentencepiece `spiece_model` where this checkpoint carries
`tokenizer_json`, and its nvfp4 support requires `comfy_kitchen` plus CUDA 13
plus cc 8.9 β the floor this file exists to avoid. The DiT only ever needed a
CONDITIONING.
> **Not yet run inside ComfyUI.** The encode path is the same code this
> project's `ltx_conditioning_dump.py` runs daily and the conditioning wrapper
> is lifted verbatim from its renderer, so the pieces are exercised β but the
> nodes themselves have not been loaded in a live ComfyUI, and that is a
> different claim.
## Try it
**[Space: LTX-2.5 Text Encoder 4bit](https://huggingface.co/spaces/topabaem/LTX-2.5-Text-Encoder-4bit-8GB-Demo)**
β text-, image- and video-to-video, running this encoder on ZeroGPU. Measured
there at 512x320, 25 frames: t2v 53.0 s, i2v 46.9 s, v2v 37.6 s.
## Image- and video-to-video
Both work, and both needed a fix ComfyUI does not ship. `LTXVAddGuide` is the
only producer of guided LTX latents and it cannot take an LTX-2.5 one: it calls
`torch.cat` on what is a `NestedTensor` for this model. The `ValueError` in that
function saying AV guides are unsupported never fires β `NestedTensor.shape`
proxies to the video half, whose channel count is exactly the 128 it checks for
β so the real failure is a `TypeError`, and the message is stale.
Everything below the node already supports AV guides: the model routes
`keyframe_idxs` to the video branch, the sampler pads a video-only denoise mask
with ones for audio, and `model_base` splits the packed mask apart again. So
`ltx_av_guide.py` unwraps the pair, runs the stock node on the video half and
re-wraps; the guide arithmetic stays the vendor's.
Measured on a 16 GB V100 at 512x320, 25 frames: t2v 98.8 s / 5.84 GiB, i2v
62.7 s / 6.80 GiB, v2v 86.5 s / 5.84 GiB. An i2v first frame lands **relL2
0.0766** from its guide image against **0.7061** for the same seed and prompt
without the guide, so the guide is honoured rather than merely accepted.
## How it was built
nvfp4 4.5 bpw (E2M1, group 16 with an fp8-e4m3 scale) on 320 projections;
int8 row-wise on `embed_tokens` and both aggregate tables; norms and asset blobs
left BF16. Per-tensor AWQ alpha search, then sequential GPTQ error compensation
(blocksize 128, percdamp 0.01), packed inside the build because the group scales
cannot be recovered afterwards.
`A3` adds three things the plain recipe lacks:
1. **No weight column is ever zeroed.** Channels are classified by an *absolute*
activation RMS, not a threshold relative to `mean(diag(H))` β the relative
rule moves with the calibration Hessian's dynamic range, which made the old
guard's best setting differ between Volta and Ada.
2. **The AWQ scale is shaped to the storage grid.** nvfp4 carries one scale per
16 input channels and its E2M1 grid spans 12:1; an unshaped per-channel scale
spans far wider inside a group and pushes the low channels under the grid
floor, where they quantize to exactly zero. Bounding the spread to 12:1 keeps
both the smoothing and the channels. The identity
`xΒ·diag(1/s) @ Q(WΒ·diag(s))α΅` holds for any `s`, so this needs no format
change.
3. **Escalating damping, recorded.** Group shaping removes conditioning that
per-channel smoothing was supplying as a side effect; `A3` needed a 10x
escalation, written into the artifact metadata so a damped build is never
silently compared with an undamped one.
## Measured limits
* **Packing is lossless.** A twin BF16 build plus a full `verify`: all 686
tensors value-exact. No drift here is attributable to the format.
* **GPTQ builds do not reproduce across GPU architectures.** Same code, plan,
calibration and guard: V100 0.06879, Ada 0.11210 on the older source. Within
one architecture they reproduce to five decimals across separate machines.
These files were built on an **A100**; quote that alongside any figure.
* **No KL or CE ratio is reported.** This artifact carries no LM head, so
vocabulary KL and CE ratio are undefined for its deployment path. Measured
instead: conditioning relL2/cosine per branch, and a five-prompt render
comparison against a BF16 conditioning from the same card, where `A3` is
closer on four of five (mean 0.05931 against `A0`'s 0.06571).
* **Not evaluated**: human listening, native DiT cross-attention KL, or whether
the remaining gap to BF16 is visible at all in finished video.
* Prompts asking for four-legged or wheeled robots still render humans. That
happens with the BF16 encoder too β a model limit, not compression damage.
## Provenance
Source `Lightricks/LTX-2.5` revision `1b92891c`+, torch 2.11.0+cu128,
transformers 5.14.1, plan `r45c`, calibration `calib-large.txt`. Evidence in
`evidence/`: gate JSON, build logs with per-layer drift and `||Q||/||W||`, the
A100 BF16 reference, all three conditionings, and fifteen rendered clips.
Earlier builds against the pre-2026-08-17 encoder, and the method results that
came from them, are at `topabaem/Pacific-LTX-2.5-Encoder-r45d`.
|