Computer Science editorial
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
The core problem
Online agent deployments accumulate execution trajectories at massive scale and behavioral diversity, for which predefined annotation criteria hardly exist. Extracting useful evidence therefore demands costly manual annotation or verifier signals that fail to scale, leaving valuable evidence buried among redundant, incomplete, and failed executions. This raises a central question: without post-execution rewards or correctness labels, how can reusable experience be distilled from the trajectories themselves?
The authors introduce **DENSE** (Distilling Evidence from Nested Subtask Executions), which organizes trajectory-derived evidence into nested shortcut trees. By consolidating redundant attempts, identifying resolved subtasks, and retaining useful steps alongside outstanding requirements, DENSE transforms noisy execution traces into structured and reusable task-solving feedback. To evaluate whether such feedback helps agents retry the same task, the authors design **REFIT**, which measures success-rate changes between the initial attempt and feedback-guided retries.
The work targets a practical gap in continual agent self-improvement: most feedback mechanisms rely on external o
Innovation
Among feedback methods without external outcome supervision, DENSE achieves the highest strict pass rate across four agent models on **Terminal-Bench 2.1**. It improves over initial attempts by **7.12–15.64 percentage points**, while using **19.0–43.6% fewer agent tokens** on retries.
On hard tasks, DENSE consistently outperforms self-reflection in cumulative pass rate across multiple feedback iterations on all four models. This indicates that the structured, evidence-grounded shortcut trees provide more reliable guidance than unstructured self-critique, particularly when the task requires sustained multi-step reasoning.
The token savings are notable: because DENSE consolidates redundant attempts and shortcuts resolved subtasks, retries avoid re-exploring dead ends, reducing the computational cost of each feedback cycle. The combination of higher pass rates and lower token usage suggests that the distilled evidence is both more accurate and more compact than raw trajectory replay or naive reflection.
Across the four evaluated agent models, the improvement range (7.12–15.64 pp) and token reduction range (19.0–43.6%) are reported as consistent, indicating that the benefit is not t
Why it matters
The central insight of DENSE is that failed, redundant, and incomplete executions are not noise to be discarded but evidence to be distilled. By organizing trajectory-derived evidence into nested shortcut trees, the method converts a liability—the accumulation of messy traces—into an asset for self-refinement.
The absence of external outcome supervision is a key differentiator. Reward models and verifiers require either human annotation or a separate learned signal, both of which struggle to scale with the behavioral diversity of online deployments. DENSE sidesteps this by mining the trajectories themselves, making it applicable in settings where correctness labels are unavailable.
The REFIT protocol provides a clean evaluation lens: it isolates the effect of feedback on retry success and token efficiency. The reported gains on Terminal-Bench 2.1, combined with the consistent advantage over self-reflection on hard tasks across multiple iterations, support the claim that structured evidence outperforms unstructured reflection for continual agent self-improvement.
Limitations and open questions remain. The paper does not specify the exact composition of the four agent models or the full task distribution of Terminal-Bench 2.1 in the provided abstract. Generalization beyond terminal-style tasks, and the interaction between shortcut tree depth and task complexity, are natural directions for future work. Nonetheless, the results position DENSE as a promising reward-free mechanism for turning execution history into reusable task-solving feedback.
Who should read this
Opening member content…