Computer Science editorial
What Makes Software Issue Resolution Tasks Difficult for Agents?
The core problem
Innovation
Task difficulty is substantially predictable from static features, achieving an AUC of 0.863. This indicates that structural properties alone can explain a large portion of agent success rates. The analysis identifies patch fragmentation and repository scale as the primary drivers of difficulty. Prompt linguistic features become visible among top contributors for tasks in the mid-band of difficulty, suggesting a layered structure where different factors dominate at different difficulty levels. The high AUC demonstrates the feasibility of pre-hoc difficulty estimation without executing the task.
A Mermaid diagram illustrating the feature extraction and prediction pipeline:
Why it matters
The findings reveal that the difficulty of an issue resolution task is encoded in its structure. Patch fragmentation—how changes are spread across files and hunks—and repository scale—size and complexity—are the most influential factors. For tasks of intermediate difficulty, prompt linguistic features (e.g., clarity, specificity) become important, indicating a layered difficulty model. This suggests that as tasks become harder, the relative importance of different feature categories shifts. The ability to predict difficulty from static features enables difficulty-controlled benchmark construction, allowing for more interpretable evaluation of agents. It also supports pre-hoc task selection and curriculum learning. The study lays the groundwork for a deeper understanding of agent limitations and for designing tasks that target specific capabilities.
A conceptual diagram of the layered difficulty structure:
Who should read this
Opening member content…