Long-form sample — an anime idol, 15 seconds
idol-15s.mp4 is the one to watch: 14.7 s, 353 frames, 512x320, 24 fps,
video and audio generated together, rendered on a 16 GB Tesla V100 from this
4-bit encoder. prompt.txt is the prompt, unedited.
It follows a lot of what was asked. Blonde hair, large blue eyes, star hair ornament, the white and light-blue frilly dress with its navy ribbon, the bright pastel bedroom with white furniture, blue curtains and a mirrored dresser. The choreography lands in order too — it opens on a hand over the lens, peeking through, then the presenting gesture, the shy eye-rub, the point at the camera.
What it does not do
It is not speaking Japanese, and probably not speaking at all. The prompt asked for it. There is audio — 14.7 s, 71% of its energy in the 300-3400 Hz band a voice occupies — but its envelope peaks at 0.20 Hz, and articulated speech modulates at 3-8 Hz, once per syllable. The structure is closer to music or a sustained vocalisation than to sentences. Judge it by ear; the measurement only says do not expect dialogue.
Style is not reliably controllable. The same prompt, same encoder, same seed, rendered at a different length or resolution, comes out photoreal instead of animated:
| file | length | size | result |
|---|---|---|---|
idol-15s.mp4 |
14.7 s | 512x320 | animated, on-prompt |
idol-10s-4bit.mp4 |
9.7 s | 512x320 | photoreal, right scene |
idol-10s-768x512.mp4 |
9.7 s | 768x512 | photoreal, wrong scene — a dark stage with an audience |
Resolution and length change the latent shape, so they change the trajectory: these are different samples, not the same sample rendered better or worse.
That is the model, not the compression
idol-10s-bf16.mp4 is the same prompt through the BF16 original, same DiT,
same seed, same size. It is photoreal too, and it tracks the 4-bit render
gesture for gesture at the same timestamps. So "2D anime animation" is being
ignored by LTX-2.5 itself at that length, and the encoder is not what dropped
it.
Frame-to-frame difference between the two, over all 233 frames: MAE 0.0563.
That is higher than the 0.0261 mean across the five benchmark prompts in
../README.md — a 317-token prompt with ten sequenced actions gives compression
more to lose than "a blacksmith hammers steel".
Sharper: the pass these clips were missing
idol-10s-1024x640-2pass.mp4 is the same prompt and seed as
idol-10s-4bit.mp4, run through the vendor's full pipeline instead of half of
it: sample at 512x320, upscale the latent 2x, sample again at 1024x640.
Left is the one-pass render enlarged afterwards, right is the two-pass render, same frame, 1:1. Eyes, teeth, the star earring and the dress frills are separate objects on the right and mush on the left — 4.1x the Laplacian variance on a common grid.
Every other clip here, and every clip in ../, is one-pass. The renderer
skipped that second pass on a written assumption that a 16 GB card could not
afford it. Measured, it peaks at 10.03 GiB and takes 312 s — it fits. So these
samples understate what the encoder drives, and the softness in them was never
the compression.
Cost
| clip | wall clock | peak VRAM |
|---|---|---|
| 15 s, 512x320 | 123 s | 6.26 GiB |
| 10 s, 512x320 | 165 s | 6.10 GiB |
| 10 s, 768x512 | 178 s | 6.53 GiB |
All on one 16 GB V100 with --fp16-unet --novram. Fifteen seconds costs no more
memory than ten; the length ceiling here was not reached.
