Past 20 seconds: what the temporal RoPE does, and what to do about it
There is a hard number, and it is in the config
LTXAVModel carries positional_embedding_max_pos=[20, 2048, 2048], and
get_fractional_positions divides each coordinate by its entry. The temporal
one is seconds, not frames — both _prepare_positional_embeddings
(model.py:841) and the AV cross-attention path (av_model.py:865) multiply
the temporal coordinate by 1 / frame_rate first.
So the temporal rotary embedding was built for 20 seconds. Past that, the fraction exceeds 1.0 and the frequencies extrapolate off the end of anything the model saw in training. This is the failure Train Short, Inference Long (arXiv 2602.14027) attributes long-horizon degradation to.
fit_temporal_positions raises max_pos[0] to the clip's own duration, mapping
it back onto [0, 1] — position interpolation, the same trick used to extend a
language model's context, and training-free. It is safe to patch at runtime
because generate_freq_grid_np's cache is keyed on theta, dimension count and
device, never on max_pos.
Measured at 25 seconds
601 frames, 1024x640, two-pass, same prompt and seed, on the Space's H200.
| segment sharpness (5 equal parts) | last/first | |
|---|---|---|
| untreated | 0.00086 0.00130 0.00143 0.00139 0.00096 | 1.12 |
| interpolated | 0.00247 0.00302 0.00286 0.00318 0.00198 | 0.80 |
The interpolated render is 2.2-2.9x sharper in every segment.
crop_ab.png at 1:1: hair strands, eyelashes and background props are resolved
on the treated side and waxy on the untreated one.
What this did not find
No progressive collapse. The premise being tested was that quality decays as the clip runs on. It does not, in either version — the last-fifth to first-fifth sharpness ratio is 1.12 untreated and 0.80 treated, and neither drifts in character, wardrobe or room across 25 seconds. Most of the long-video literature found in this search — compounding error, history forgetting, attribute drift — describes autoregressive or sliding-window generation, where each window conditions on the last. LTX-2.5 as driven here samples the whole clip at once, so that class of failure does not apply and searching for it would have been looking in the wrong place.
What extrapolation costs here is detail across the whole clip, not stability at the end.
Limits
- The two renders are different trajectories, not the same trajectory rendered
twice, so the sharpness gap is what changed and not a controlled isolation of
cause. Same prompt, seed, size, sampler and refine strength; only
max_pos[0]differs. - One duration (25 s), one prompt, one seed. 30 s, 40 s and beyond were not tried, and the point where interpolation itself starts to hurt — it compresses temporal detail by construction — was not located.
- Below 20 s this changes nothing and is not applied.
- Laplacian variance tracks contrast as well as sharpness. It is used here between two renders of the same scene, which is the only place it is fair.
Files
idol25-fit.mp4, idol25-nofit.mp4, ab.png (timeline), crop_ab.png (1:1).