FocusVTC

Paper | Code

FocusVTC is a 9B vision-language model based on Qwen3.5. It is optimized for long-context visual document understanding and evidence-focused tool use. The model can reason over document-page images, decide when a higher-resolution crop is useful, and continue answering after receiving the cropped region as a new visual observation.

Quick start

FocusVTC requires a recent Transformers version with Qwen3.5 support.

pip install "transformers>=5.16.1" accelerate pillow
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "zfz04/FocusVTC"

processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    dtype="auto",
    device_map="auto",
)

image = Image.open("document_page.png").convert("RGB")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image"},
            {"type": "text", "text": "Answer the question using the document."},
        ],
    }
]

prompt = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
inputs = processor(
    text=[prompt],
    images=[image],
    return_tensors="pt",
).to(model.device)

output_ids = model.generate(**inputs, max_new_tokens=512)
new_tokens = output_ids[:, inputs.input_ids.shape[1]:]
answer = processor.batch_decode(new_tokens, skip_special_tokens=True)[0]
print(answer)

For multi-page documents, add one image placeholder per page and pass the page images to the processor in the same order.

Visual tool use

FocusVTC is designed to work in an agent loop with the following tool:

{
  "type": "function",
  "function": {
    "name": "zoom_region",
    "description": "Crop a potentially unreadable region from a document page and return it as a new image.",
    "parameters": {
      "type": "object",
      "properties": {
        "page": {
          "type": "integer",
          "description": "1-based page number."
        },
        "bbox_2d": {
          "type": "array",
          "items": {"type": "number"},
          "minItems": 4,
          "maxItems": 4,
          "description": "[x1, y1, x2, y2] normalized to [0, 1000]."
        }
      },
      "required": ["page", "bbox_2d"]
    }
  }
}

The host application is responsible for executing the crop on the corresponding high-resolution page and appending the result as a new visual observation. Plain Transformers generation does not execute the tool automatically.

Intended use

  • Long and multi-page document question answering
  • Fine-grained visual evidence localization
  • Visual retrieval and inspection with iterative region zooming
  • Research on multimodal agents and visual tool use

Limitations

  • The model may produce incorrect answers or inaccurate crop coordinates.
  • Reliable tool use requires an external runtime that validates and executes zoom_region calls.
  • Performance can depend strongly on page resolution, prompt format, generation settings, and the amount of visual context.
  • Outputs should be independently verified in high-stakes settings.

Citation

If you find FocusVTC useful in your research, please cite our paper:

@misc{zhong2026focusvtcefficienthighperformancevisual,
      title={FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution},
      author={FangZhi Zhong and Xuerui Qiu and Yuqi Pan and Ya Liu and Shaowei Gu and Bo Xu and Guoqi Li},
      year={2026},
      eprint={2609.36651},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.36651},
}
Downloads last month
2
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zfz04/FocusVTC

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(968)
this model

Paper for zfz04/FocusVTC