FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
Abstract
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at 2.9times input compression, including tool observations, versus 57.5 for Glyph at 3.0times input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a 2.79times online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
Community
FocusVTC reads long documents through compact page images and retrieves higher-resolution evidence as it reasons. Built around Qwen3.5, it combines Reasoning–Evidence Localization supervised fine-tuning (REL-SFT) with tool-assisted GRPO. The model uses zoom_region to inspect a selected region of an aligned high-DPI page before answering.
Figure 1 from the paper. Fixed-resolution VTC and FocusVTC's adaptive reading strategy, with RULER v1 and general-capability comparisons.
Key Features
- Adaptive resolution. Low-DPI pages provide the document overview; selected regions are read from aligned high-resolution pages.
- Evidence-grounded supervision. REL-CoT connects reasoning and answers with evidence page numbers and bounding boxes.
- Tool-assisted learning. GRPO trains the policy to request and use evidence crops, with rewards for answer accuracy and evidence-aware tool use.
- Training and evaluation code. The release includes REL-CoT preparation, SFT and GRPO runtimes, document inference, and benchmark adapters.
Get this paper in your agent:
hf papers read 2609.36651 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
zfz04/REL-CoT
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper

