What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents
Abstract
As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications. We introduce AgentMonBench, a software-engineering benchmark comprising three subsets that cover two complementary dimensions: alignment between requirements and behavior, and awareness of consequential autonomous decisions for verification. To support these judgments, we propose the Evidence-Grounded Behavior Graph (EBG), a training-free method that groups source-linked evidence into behaviors and organizes their relationships into a graph. EBG presents task-oriented views of this graph to help monitors interpret behavior in context. Experiments across eight models show that EBG improves decision identification and evidence localization in most settings compared with direct access to the original context. Further experiments show that EBG's evidence-localization gains persist across input scales and hyperparameter settings, while real-world applications illustrate its practical value for human oversight.
Community
As AI Agents Become More Autonomous, How Can We Know What They Actually Did?
As AI agents become increasingly capable, the tasks they undertake are evolving from simple question answering and tool use to autonomous workflows that may last for hours or even days. From software development and data analysis to scientific research, agents can independently plan, invoke tools, modify code, and continuously adjust their strategies based on intermediate results.
Traditionally, research on AI interpretability has focused primarily on understanding the internal decision-making mechanisms of models. However, as agents take on increasingly complex, long-horizon tasks, a new challenge is emerging: not only are the internal mechanisms of AI models difficult to interpret, but their long-horizon behaviors are also becoming increasingly difficult to oversee.
Giving an agent a goal does not mean that every constraint and decision along the way has been fully specified. For example:
- Underspecified requirements: When user instructions are incomplete, agents may autonomously fill in missing assumptions, which may not align with the user's actual intentions.
- Silent behavioral deviations: Agents may modify their original plans or implementations without informing users. Even when all tests pass, their actual behavior may still deviate from the intended requirements.
- Unverified consequential decisions: Seemingly reasonable decisions, such as changing evaluation metrics, filtering data, or adjusting experimental settings, may fundamentally affect the final outcome without ever being confirmed by the user.
More importantly, as execution trajectories grow longer, critical evidence becomes scattered across extensive conversations, tool calls, and intermediate artifacts. Users cannot realistically inspect every step. Sometimes, they may not even know which consequential decisions the agent has made, when intervention is necessary, or what follow-up instructions they should provide.
This creates a fundamental challenge for traditional Human-in-the-Loop paradigms: humans may remain part of the oversight process without genuinely understanding what the agent has actually done.
Motivated by this problem, in our recent work, "What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents," we explore a new direction:
Can we build an independent, evidence-grounded oversight system that proactively identifies decisions worth human attention and verification, rather than requiring users to manually search through lengthy execution logs?
We investigate this problem from two complementary perspectives:
① Alignment between Requirements and Behavior
Does the agent's actual behavior align with the user's requirements? When user specifications leave important details unresolved, or when an agent silently deviates from the intended behavior, can an oversight system identify issues that warrant clarification or verification?
② Awareness and Verification of Consequential Decisions
What consequential decisions has the agent made during autonomous execution? Even when these decisions are not necessarily incorrect, can an oversight system proactively surface them and provide sufficient evidence for users to assess whether they should be accepted?
Building on these two dimensions, we introduce AgentMonBench, a benchmark comprising three subsets—SpecGAP, SilentSwap, and FeedbackTrace—to systematically evaluate the ability of LLMs to identify consequential decisions and localize supporting evidence.
We further propose the Evidence-Grounded Behavior Graph (EBG), which organizes evidence scattered across repositories and execution trajectories into a structured behavioral graph. This enables oversight models to interpret an agent's actual behavior in the context of task requirements, identify consequential decisions, and trace supporting evidence back to its original sources.
Importantly, EBG is training-free. We also integrate it into the Codex Harness, introducing oversight checkpoints before execution, after consequential execution changes, and before final reporting. These checkpoints help users clarify requirements, review important behavioral changes, and verify outcomes based on evidence.
Experiments across eight models show that EBG improves consequential decision identification and evidence localization in most evaluation settings. Case studies involving real research workflows further demonstrate its potential to support user verification and guide subsequent agent execution.
More broadly, we hope this work encourages a shift in how we think about AI interpretability and oversight.
Interpretability should not be limited to answering "Why does a model reason this way internally?" It should also address "What did an agent actually do throughout a long-horizon task? Which decisions shaped the outcome? And what evidence do we have to trust those decisions?"
As AI agents become increasingly autonomous, effective human oversight may no longer mean participating in every step. Instead, it may mean empowering humans to understand, verify, and intervene at the moments that truly matter, with reliable evidence.
From model-level interpretability to long-horizon behavioral oversight. From Human-in-the-Loop to meaningful Human Oversight.
📄 Paper: https://arxiv.org/abs/2610.06406
💻 Code: https://github.com/zhk-lab/EBG
📊 Benchmark: https://huggingface.co/datasets/ZhaoHongKang/AgentMonBench
🌐 Project Page: https://zhk-lab.github.io/agentmonbench-ebg/
This is the question I care about most. In my own system every value carries who asserted it and how strong the evidence is, so a person only has to look closely at the decisions that matter. I've only read the abstract, so forgive me if this is covered. Does EBG keep a link back to the exact source line, so a reviewer can check the evidence instead of trusting the summary?
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper