Introduces an online DAG-replanning algorithm for deep-research agents that repairs task graphs incrementally as evidence emerges rather than only after failures.
DAGent introduces an Evaluate-then-Grow planning paradigm that incrementally expands a task graph based on confidence and uncertainty signals from completed nodes, rather than committing to a full plan upfront. This approach replaces the brittle Plan-then-Patch strategy with an orchestrator that grows the DAG one batch at a time, conditioning each expansion on structured per-node feedback.
Key features include: Incremental Expansion: The system appends new nodes only when evidence supports it, avoiding wasted computation on invalid branches. Hierarchical Context: Uses compact QueryDocs for default propagation and full InteractionTranscripts for on-demand recall. * DAGRPO: A reinforcement learning adaptation that injects topology-conditioned credit based on the recorded DAG structure.
DAGent outperforms strong open-source baselines by 5.3 points on BrowseComp-Plus and 5.8 points on GAIA at the Qwen3-235B-A22B scale.
DAGent addresses a central weakness in deep-research agents: they often build a long-horizon plan up front, execute it, and only recover when a downstream step fails or produces poor results. The paper proposes an online DAG-replanning framework in which the agent’s plan is represented as a directed acyclic graph of subtasks, evidence requirements, and evaluation checkpoints. Rather than treating this graph as static, the agent executes partial subgraphs, evaluates the accumulated evidence, and then incrementally grows or repairs the graph by adding, removing, or reordering nodes. This “evaluate-then-grow” loop turns evidence into a planning signal: gaps, contradictions, or weak support discovered early can trigger local revisions before the system has invested in an entire brittle trajectory.
Methodologically, the contribution is a control algorithm for long-horizon information-seeking under partial observability. The key insight is that deep research is not just task decomposition; it is also evidence management. A subtask may be syntactically well-formed but epistemically premature, while a seemingly minor observation may invalidate a branch or reveal a more efficient decomposition. By embedding evaluation gates inside the DAG, DAGent can distinguish between operational failure and informational insufficiency, and can repair the plan incrementally instead of restarting from the root or waiting for terminal failure. This should reduce wasted search, improve robustness to ambiguous or nonstationary evidence, and make agent trajectories more interpretable because each revision is tied to explicit evaluation outcomes.
The work matters because modern research agents are moving from single-shot question answering to multi-step investigations that require planning, tool use, and synthesis over large evidence spaces. Static plan-then-execute approaches struggle when the right next action depends on what was learned a few steps earlier. DAGent offers a more adaptive architecture for such settings, potentially improving sample efficiency and reliability on open-ended research tasks. More broadly, it points toward a design principle for agentic systems: plans should be treated as living hypotheses that are continuously evaluated and reshaped by the evidence they produce.