arxiv:2606.02320

TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation

Published on Jun 1

· Submitted by

Xinkai Ma on Jun 2

NJU-LINK Lab

Upvote

Authors:

Xinkai Ma ,

Abstract

A multimodal deep research benchmark and agent framework are introduced to evaluate and improve the factual reliability and visual alignment of automated report generation systems.

Generated by Qwen/Qwen2.5-Coder-32B-Instruct

Deep Research Agents have shown strong capability in multi-step information retrieval, reasoning, and long-form report generation, but existing benchmarks and systems remain predominantly text-centric, with limited evaluation of whether visual elements are factually reliable and well aligned with the surrounding analysis. To address this gap, we introduce TVIR (Text--Visual Interleaved Report Generation), which includes TVIR-Bench, a benchmark of 100 expert-curated multimodal deep research tasks that require visual elements to serve specific analytical sub-goals, and TVIR-Agent, a hierarchical multi-agent framework that serves as a strong baseline for constructing outlines, retrieving images, generating charts with traceable sources, and composing reports through context-aware sequential writing. We further develop a dual-path evaluation framework that combines Textual Assessment and Visual Assessment. Experiments across nine deep research systems show that TVIR-Agent achieves strong overall performance, underscoring the importance of explicit multimodal design and evaluation for evidence-driven report generation.

View arXiv page View PDF Project page GitHub 3 Add to collection

Community

Cenji630

Paper author Paper submitter about 7 hours ago

We present TVIR, the first benchmark and agent framework specifically designed for text-visual interleaved report generation. Unlike existing text-only deep research systems, TVIR-Bench evaluates both textual quality and visual integration across 100 expert-curated tasks. Our TVIR-Agent achieves state-of-the-art performance, demonstrating that structured multi-agent collaboration is key to generating high-quality multimodal reports.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Upvote

Get this paper in your agent:

hf papers read 2606.02320

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.02320 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.02320 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.