Papers
arxiv:2609.37236

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

Published on Sep 29
· Submitted by
Ido Levy
on Sep 30
Authors:
,
,
,

Abstract

An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model 15times larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the 15times larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.

Community

Paper author Paper submitter

What should an LLM agent pursue that the user never asked for? We define horizontal proactivity (a need the current state already names) and vertical proactivity (a need only newly found evidence names), measure both against need graphs with no LLM judge, and train a questioner with Q&D, which prefers the question whose continuation retrieves more of the required evidence.

On MuSiQue at equal retrieval spend, the trained Qwen3-8B questioner recovers 90% of the required evidence vs 78% for the same model prompted, and it outperforms GPT-OSS-120B (15× larger) in the same role on 2 of 3 benchmarks. Without further training, it more than doubles retail task success in a τ²-bench customer-service agent.

Code, metrics and the trained model are open.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.37236
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 3

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.37236 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.37236 in a Space README.md to link it from this page.

Collections including this paper 1