Papers
arxiv:2609.33659

Learning Multimodal Embeddings with Evidence-Aligned Readout

Published on Sep 27
· Submitted by
zrchen
on Oct 8
Authors:
,
,
,
,
,
,
,
,
,

Abstract

Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled 2times3 study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.

Community

Paper submitter

Excited to share our work EviAlign: Learning Multimodal Embeddings with Evidence-Aligned Readout!
Can the semantic structure of generated evidence guide how multimodal representations are extracted?
We introduce EviAlign, a framework that jointly designs semantic evidence generation and boundary-aware embedding readout within a unified MLLM.
🔹 Evidence-Aligned Readout: Extract representations at semantic evidence boundaries rather than relying on a single final token.
🔹 Generation–Retrieval Co-training: Jointly optimize evidence generation and contrastive retrieval.
🔹 Strong Retrieval Performance: Achieve 76.9 average Recall@1 across 12 MMEB tasks using 500K training pairs, while maintaining efficient single-vector retrieval.
Our findings highlight the importance of co-designing evidence organization and representation readout for multimodal retrieval.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33659
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33659 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33659 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33659 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.