Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents
Abstract
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model 15times larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the 15times larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.
Community
What should an LLM agent pursue that the user never asked for? We define horizontal proactivity (a need the current state already names) and vertical proactivity (a need only newly found evidence names), measure both against need graphs with no LLM judge, and train a questioner with Q&D, which prefers the question whose continuation retrieves more of the required evidence.
On MuSiQue at equal retrieval spend, the trained Qwen3-8B questioner recovers 90% of the required evidence vs 78% for the same model prompted, and it outperforms GPT-OSS-120B (15× larger) in the same role on 2 of 3 benchmarks. Without further training, it more than doubles retail task success in a τ²-bench customer-service agent.
Code, metrics and the trained model are open.
Get this paper in your agent:
hf papers read 2609.37236 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 3
dolev31/ProactiveInquirer-Qwen3-8B
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper