Papers
arxiv:2609.14302

E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning

Published on Sep 13
· Submitted by
wangjunjie
on Sep 15
Authors:
,
,

Abstract

A benchmark for financial chart reasoning reveals that vision-language models often fail to maintain traceable evidence-to-action reliability, highlighting the need to evaluate the full reasoning chain rather than single hallucination scores.

Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not whether evidence remains traceable through rationale, confidence, and final action. We introduce E2A-Bench, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors. E2A-Bench evaluates grounding, reasoning-action consistency, evidence-confidence calibration, and directional coverage through UCR, RCI, ECI, and NDR, where NDR measures coverage-aware evidence-to-action reliability rather than realized trading performance. Evaluating 20 VLMs reveals three failures hidden by scalar hallucination scores: the lowest-UCR model ranks near the bottom by NDR due to only 6.4% directional coverage; oracle-aided verification reduces unsupported claims but can collapse coverage; and financial fine-tuning amplifies the BUY:SELL ratio by factors of 4.21 to 4.68 across strict base-fine-tuned pairs. These results show that financial VLM evaluation should trace the full evidence-to-action chain rather than rely on a single hallucination score. Code and data: https://github.com/wanng-ide/E2A-Bench

Community

Paper author Paper submitter

E2A-Bench is a multimodal benchmark for evaluating evidence-to-action reliability in financial chart reasoning. It assesses whether VLMs can ground their analysis in chart evidence and maintain consistency across evidence, reasoning, confidence, and final decisions.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.14302
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.14302 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.14302 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.14302 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.