Papers
arxiv:2609.36651

FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

Published on Sep 29
· Submitted by
Zhong FangZhi
on Sep 30
Authors:
,
,
,
,
,

Abstract

Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at 2.9times input compression, including tool observations, versus 57.5 for Glyph at 3.0times input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a 2.79times online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.

Community

Paper author Paper submitter
•
edited about 4 hours ago

FocusVTC reads long documents through compact page images and retrieves higher-resolution evidence as it reasons. Built around Qwen3.5, it combines Reasoning–Evidence Localization supervised fine-tuning (REL-SFT) with tool-assisted GRPO. The model uses zoom_region to inspect a selected region of an aligned high-DPI page before answering.

屏幕截图 2026-09-30 222212

Figure 1 from the paper. Fixed-resolution VTC and FocusVTC's adaptive reading strategy, with RULER v1 and general-capability comparisons.

Key Features

  • Adaptive resolution. Low-DPI pages provide the document overview; selected regions are read from aligned high-resolution pages.
  • Evidence-grounded supervision. REL-CoT connects reasoning and answers with evidence page numbers and bounding boxes.
  • Tool-assisted learning. GRPO trains the policy to request and use evidence crops, with rewards for answer accuracy and evidence-aware tool use.
  • Training and evaluation code. The release includes REL-CoT preparation, SFT and GRPO runtimes, document inference, and benchmark adapters.

屏幕截图 2026-09-30 222220

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.36651
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36651 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.