Ilmu Komputer & AI editorial
Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark
The core problem
Innovation
The audit leverages public SWE-bench Verified trajectories and a frozen dual-head early-outcome prediction pipeline, which outputs separate predictions for SUCCESS and FAILURE outcomes. The primary analysis is a leave-one-agent-out calibration audit: for each target agent, the predictor is trained on all other agents and evaluated on the held-out agent. To control for shared predictor effects, a leave-two-agents-out control is also performed. Oracle prior correction adjusts for base rate differences. A robustness battery tests the stability of findings across variations in training cohorts, task resampling, task halves, jackknife resampling, and decision thresholds. The pre-registered heterogeneity criterion assesses whether pairwise corrected-gap differences are broadly distributed. Additionally, a fixed-scaffold TerminalBench analysis serves as a boundary test for cross-benchmark replication. The key metric is the corrected gap, defined as the difference between predicted confidence and observed accuracy, with calibration-transfer error indicated by persistent nonzero gaps. Formally, for a target agent and head
Why it matters
Who should read this
Opening member contentโฆ