Papers
arxiv:2610.04318

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

Published on Oct 3
ยท Submitted by
Sixun Dong
on Oct 6
Authors:
,
,
,
,
,
,

Abstract

Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding latency depends on the size of the candidate pool rather than the final token budget. An empirical study across multiple VLMs and long-video benchmarks yields three findings: dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets; resolution-sensitive tasks benefit from selected high-resolution frames; and front-end decoding dominates wall time for hour-long videos. Motivated by these findings, we introduce LoHi, a training-free, single-pass framework that combines a dense low-resolution video stream with sparse high-resolution image frames through the VLM's native video and image pathways. LoHi-Anchor selects high-resolution frames using codec I-frame metadata, while LoHi-SemDiv uses query relevance and visual diversity over CLIP features. Across three long-video benchmarks, LoHi improves average accuracy by 10.6 percentage points over the native-resolution baseline at a matched token budget and by 5.2 percentage points over the strongest prior efficiency method. It also reduces front-end decoding latency by up to 7x on hour-long videos. Project page: https://sixundong.com/projects/lohi

Community

Paper submitter

๐ŸŽฌ LoHi (NeurIPS 2026): Rethinking long-video efficiency.

๐Ÿ’ก Three lessons

  1. ๐ŸŽž๏ธ More frames, not more pixels: at the same token budget, dense low-resolution frames beat sparse native-resolution frames.
  2. ๐Ÿ” Resolution is task-dependent: most questions are fine at low resolution; OCR and fine details need high resolution.
  3. โฑ๏ธ Decoding matters: on long videos, frame decoding can take longer than the model itself.

โš ๏ธ Limits of keyframe selection & token pruning: both stay at native resolution and discard afterward. They decode 256 full frames only to keep 16 (or drop ~94% of tokens), paying the decoding cost while losing temporal coverage. A plain low-resolution baseline beats both (64.4 vs. 60.8 / 59.3 on VideoMME).

๐Ÿš€ Our simple method: LoHi decodes 128 frames once, feeds dense low-resolution video plus a few high-resolution images. Training-free, single pass. 66.8 on VideoMME with 3.9s time-to-first-token (vs. 6.0โ€“9.0s) on Qwen3-VL-4B.

๐Ÿ”— Project page with interactive demos for each lesson (try the resolution dial and the decoding cost of any frame): https://sixundong.com/projects/lohi

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.04318
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.04318 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.04318 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.04318 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.