Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

What Makes Software Issue Resolution Tasks Difficult for Agents?

A measurement framework reveals that task difficulty is predictable from static structural properties, driven by patch fragmentation and repository scale.
Ebtesam Al-Haque; Brittany Johnson· 2026· DOI 10.48550/arXiv.2608.18280

The core problem

Agentic systems are rapidly saturating benchmarks, yet benchmark scores remain difficult to interpret due to a lack of control and characterization of task difficulty. We currently have little understanding of what makes one software issue resolution task harder than another, and to what extent difficulty is predictable from static task properties. This paper addresses that gap by proposing a measurement framework to systematically quantify which structural properties of software tasks correspond to agent success rates. The authors conduct a large-scale empirical study on CoderForge-Preview, the largest open dataset of coding agent trajectories to date, extracting features across task patch, repository, and prompt. The goal is to enable static, pre-hoc difficulty estimation and lay the groundwork for difficulty-controlled benchmark construction.

Innovation

Task difficulty is substantially predictable from static features, achieving an AUC of 0.863. This indicates that structural properties alone can explain a large portion of agent success rates. The analysis identifies patch fragmentation and repository scale as the primary drivers of difficulty. Prompt linguistic features become visible among top contributors for tasks in the mid-band of difficulty, suggesting a layered structure where different factors dominate at different difficulty levels. The high AUC demonstrates the feasibility of pre-hoc difficulty estimation without executing the task.

A Mermaid diagram illustrating the feature extraction and prediction pipeline:

Agentic systems are rapidly saturating benchmarks, yet benchmark scores remain difficult to interpret due to a lack of control and characterization of task difficulty. We currently have little understanding of what makes one software issue resolution task harder than another, and to what extent difficulty is predictable from static task properties. This paper addresses that gap by proposing a measurement framework to systematically quantify which structural properties of software tasks correspond to agent success rates. The authors conduct a large-scale empirical study on CoderForge-Preview, the largest open dataset of coding agent trajectories to date, extracting features across task patch, repository, and prompt. The goal is to enable static, pre-hoc difficulty estimation and lay the groundwork for difficulty-controlled benchmark construction.
The study analyzes CoderForge-Preview, a dataset of coding agent trajectories. Features are extracted from three dimensions: task patch (e.g., lines changed, files touched, fragmentation), repository (e.g., size, complexity, number of contributors), and prompt (e.g., linguistic properties such as readability, specificity). The predictive power of each feature against task outcomes is evaluated using ensemble methods, SHAP attribution, and effect size analysis. The target outcome is agent success rate on issue resolution tasks. The framework aims to quantify the relationship between static task properties and agent performance, enabling pre-hoc difficulty estimation.

Why it matters

The findings reveal that the difficulty of an issue resolution task is encoded in its structure. Patch fragmentation—how changes are spread across files and hunks—and repository scale—size and complexity—are the most influential factors. For tasks of intermediate difficulty, prompt linguistic features (e.g., clarity, specificity) become important, indicating a layered difficulty model. This suggests that as tasks become harder, the relative importance of different feature categories shifts. The ability to predict difficulty from static features enables difficulty-controlled benchmark construction, allowing for more interpretable evaluation of agents. It also supports pre-hoc task selection and curriculum learning. The study lays the groundwork for a deeper understanding of agent limitations and for designing tasks that target specific capabilities.

A conceptual diagram of the layered difficulty structure:

Who should read this

CS practitioners and researchers

Opening member content…