Abstract
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent
Community
Hi everyone! Thanks for checking out our paper (NeurIPS 2026). ๐
Why should a deep research agent commit to its whole task graph before it has seen any evidence?
We introduce DAGent, a DAG-based multi-agent framework that replaces Plan-then-Patch planning (a full task DAG committed before execution and repaired afterwards) with Evaluate-then-Grow: an Orchestrator grows the task DAG one batch at a time, conditioning each expansion on the confidence and uncertainty reported by completed nodes. The append-only DAG also defines structural RL signals, which DAGRPO uses for topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans.
On BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and DAGRPO improves over GRPO with the same agent and training budget by 3.0 average Pass@1 points at the Qwen3-8B scale.
One limitation: the RL experiments train Qwen3-8B with LoRA for 21 update steps, so whether the DAGRPO gains persist at larger training scales, longer schedules, or full fine-tuning is not established by this paper.
We'd love to hear your thoughts: when should a planner commit to structure, and when should it wait for evidence? Happy to discuss the method and experiments!
Get this paper in your agent:
hf papers read 2609.39154 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper