topabaem's picture
Replace the samples with a BF16 comparison set
77ec797 verified
|
Raw
History Blame Contribute Delete
7.54 kB
# Samples β€” 4-bit against the BF16 original
Every clip here exists twice: once from the **BF16 original (26.264 GB)** and
once from this **4-bit build (8.46 GB)**. `compare-NN.mp4` puts the pair side by
side, BF16 on the left.
Everything downstream of the text encoder is held identical β€” same DiT
(`LTX25-distilled-DiT-Q3_K_M`), same seed, same 8-step distilled schedule, same
VAE tiling, one process, one GPU. The encoder is the only thing that differs.
768x512, 24 fps, 2.0 s, video and audio generated together, rendered on a
**Tesla V100-SXM2-16GB** β€” the card this compression exists to fit.
## Two sets, because they answer different questions
`compare-NN.mp4` uses **`euler_ancestral`**, the vendor default and what you
actually get. It injects fresh noise at every step, so a conditioning difference
of cosine 0.998 does not stay small β€” it steers the trajectory elsewhere and the
two builds return different **takes of the same scene**. Look at these to see
what the model produces. They cannot tell you what the compression cost, and
reading them that way will mislead you: in `compare-01` the BF16 side has a
second figure on the dune that the 4-bit side does not, and that is the sampler
re-rolling, not the encoder forgetting.
`compare-det-NN.mp4` uses **`euler`**, which injects nothing. Same seed means
the same starting noise, so the only thing left that can move a pixel is the
conditioning. **This is the set that attributes a difference to the encoder.**
It is worth saying what this did *not* do. Removing the ancestral noise did not
reliably bring the pairs closer β€” measured below, it helped enormously on one
clip and made three marginally worse. Fixing the starting noise does not fix
the path: eight distilled steps through a 22B transformer are nonlinear enough
that a cosine-0.998 conditioning difference can still land somewhere else. So
the deterministic set is the honest attribution, not a flattering one.
## Prompts
Written to pull on what a text encoder has to keep apart: camera instruction,
shot scale, lighting direction, surface material, and a distinct audio scene.
| # | seed | prompt |
|---|---:|---|
| 00 | 20260813 | A humanoid robot steps out of a freight elevator into a rain-slick loading bay, sodium lights flaring off its wet white shell. Locked-off low camera as the robot turns toward the lens. Audio: heavy rain on metal, one servo whine, distant thunder. |
| 01 | 20260814 | A lone astronaut in a scuffed white suit walks across a red dune at sunset, a long shadow stretching behind, dust curling off each bootfall. Low tracking camera moving with the walk. Audio: wind over sand, suit fan hum, muffled breathing. |
| 02 | 20260815 | A blacksmith hammers glowing orange steel on an anvil inside a dark forge, sparks bursting outward with every strike, warm rim light on soot-streaked arms. Medium shot, camera static. Audio: ringing hammer strikes, roaring bellows, a hiss as the blade is quenched. |
| 03 | 20260816 | A vintage sports car slides through a wet mountain hairpin at night, headlights sweeping across the guardrail, spray lit red by the tail lights. Camera pans to follow the car through the turn. Audio: engine roar, tyre squeal, rain on asphalt. |
| 04 | 20260817 | A battered rescue robot pushes through thick smoke in a collapsed corridor, headlamp cutting a cone through the dust, orange emergency light pulsing on the walls. Handheld camera following from behind. Audio: crackling fire, a muffled alarm, grinding treads. |
## Drift, measured before any of it reaches the DiT
4-bit conditioning against BF16 conditioning, same prompts, same tokenizer path:
| # | video relL2 | audio relL2 | video cosine | audio cosine |
|---|---:|---:|---:|---:|
| 00 | 0.05820 | 0.05010 | 0.99831 | 0.99875 |
| 01 | 0.05614 | 0.04925 | 0.99843 | 0.99879 |
| 02 | 0.07002 | 0.06274 | 0.99755 | 0.99803 |
| 03 | 0.05959 | 0.05222 | 0.99823 | 0.99864 |
| 04 | 0.05203 | 0.04624 | 0.99865 | 0.99893 |
| **mean** | **0.05919** | **0.05211** | | |
Prompt 02, the forge, is hardest on both branches. It also has the most
simultaneous physical events β€” sparks, glow, bellows, quench β€” which is where a
text encoder has the most to lose.
## Difference on screen
Mean absolute error per channel over every decoded frame, 4-bit against BF16.
Not relative L2: these scenes run from a night forge to a sunlit dune, and
relative error divides by the frame's own norm, so a dark clip would score worse
than a bright one for the same visible difference. MAE has its own bias β€” it
flatters dark frames β€” so **only the ratio compares across rows**, because
within one scene the brightness cancels.
| # | scene | `euler_ancestral` | `euler` | ratio |
|---|---|---:|---:|---:|
| 00 | robot, rain | 0.0172 | 0.0182 | 0.94x |
| 01 | dune | 0.0596 | 0.0580 | 1.03x |
| 02 | forge | 0.0098 | 0.0129 | 0.76x |
| 03 | night road | 0.0702 | 0.0252 | **2.79x** |
| 04 | smoke | 0.0128 | 0.0165 | 0.77x |
| | **mean** | **0.0339** | **0.0261** | 1.26x |
## The one that actually diverged
Under deterministic sampling, four of the five pairs hold together: same
composition, same lighting, same timing, differing in surface detail. Clip 00
gives the 4-bit side slightly more panel detail on the hull; clips 02, 03 and 04
are close enough that finding the difference takes looking.
**Clip 01 is not.** Same terrain, same sun, same shadow, same walk β€” and BF16
renders a soldier in helmet and fatigues while the 4-bit renders a man in a
business suit. The sampler is deterministic and the seed is identical, so that
belongs to the encoder.
It deserves one qualification rather than being either buried or oversold.
Neither build followed the prompt: it asked for an astronaut in a scuffed white
suit and got neither. The model has no confident answer here, so it lands on an
arbitrary human, and in that regime a small conditioning change flips which one.
The difference is real and attributable; its size reflects a prompt that was
already unstable on both sides.
## What these prove, and what they do not
They show that at 8.46 GB the encoder drives the same DiT to a comparable result
on five prompts, at one resolution, one length and one seed each.
They are **not** a claim that the 4-bit encoder is indistinguishable from BF16.
No blind study, no listening test, no sweep over seeds or resolutions.
**Prompt adherence is partial, and mostly on both builds equally.** Clip 00 has
no freight elevator and frames far tighter than "locked-off low camera" implies;
clip 04 reads closer to a firefighter than a rescue robot; clip 01 produces no
astronaut on either side. The BF16 original does the same thing on the same
prompts, so that is a limit of LTX-2.5 distilled rather than compression damage.
The clips are published unedited β€” including clip 01, where the compression
does visibly cost something β€” rather than replaced with prompts that happen to
flatter the result.
## Files
* `compare-NN.mp4` β€” BF16 left, 4-bit right, `euler_ancestral`.
* `compare-det-NN.mp4` β€” the same pairing under deterministic `euler`.
* `bf16-NN.mp4`, `4bit-NN.mp4` β€” the individual `euler_ancestral` renders.
* `bf16-det-NN.mp4`, `4bit-det-NN.mp4` β€” the individual `euler` renders, if you
want to step through them frame by frame rather than trust the stack.
Audio in the paired clips is taken from the 4-bit side rather than mixed: the
two tracks are different renders of the same prompt, and mixing them would
produce a phasing artefact belonging to neither build.