Computer Science editorial
REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring
The core problem
Large Language Models (LLMs) offer new opportunities for automated code refactoring, yet generated changes must reduce targeted quality problems without introducing new issues or altering behaviour-relevant code structures. This dual requirement—effective smell reduction coupled with behavioural preservation—defines the central challenge addressed by REFINE (Refactoring with Evidence-aware Flow for Integrated ageNtic Execution).
The authors, Muhammad Waseem, Aakash Ahmad, and Pekka Abrahamsson, position REFINE as a tool-agnostic, evidence-aware multi-agent approach for generating Java file-level refactoring candidates. Rather than relying on a single monolithic LLM prompt, REFINE orchestrates a pipeline that combines static-analysis-guided smell identification, smell-informed planning, LLM-based transformation, automated re-analysis, preservation checks, and structured reporting. The core research question is whether an evidence-guided, multi-agent decomposition of the refactoring workflow can outperform a direct-prompt baseline in terms of code-smell reduction, edit size, and preservation of public APIs.
The evaluation spans 450 Java files drawn from 15 open-source systems, prod
Innovation
REFINE reduces detected code smells by 68.26%, 72.79%, and 68.49% across the OpenAI GPT-5.5, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.8 configurations, respectively. The strongest reductions are observed for major smells, indicating that the evidence-guided pipeline is most effective on higher-severity quality problems. Across the three configurations, the reduction rates cluster in a relatively narrow band of roughly 68–73%, suggesting that the multi-agent structure provides consistent benefits across different underlying models.
The matched 150-file direct-prompt baseline shows that REFINE achieves a higher median code-smell reduction with smaller edits and fewer public-method removals. This comparison is central to the paper's contribution: the evidence-aware, multi-agent decomposition outperforms a direct-prompt approach not only on the primary smell-reduction metric but also on edit economy and API preservation. Smaller edits imply less churn and lower review burden, while fewer public-method removals imply reduced risk of breaking downstream consumers.
However, the results are not uniformly positive. Broader quality improvements are inconsistent, meaning t
Why it matters
The REFINE results support a nuanced conclusion: evidence-guided, multi-agent LLM refactoring can substantially reduce detected code smells and can outperform direct prompting on median smell reduction, edit size, and public-method preservation, but it does not eliminate the need for downstream verification. The 68.26%–72.79% smell-reduction range across three distinct LLM configurations indicates that the pipeline's benefits are not tied to a single model vendor, which strengthens the tool-agnostic claim. The observation that major smells are reduced most strongly suggests the pipeline prioritizes high-impact quality problems, consistent with its smell-informed planning stage.
The persistence of residual risks—assert/fail-call changes and public-method removal—is the most consequential finding for practice. These are precisely the kinds of changes that can alter behaviour or break dependent code, and their detection by preservation checks demonstrates the value of the evidence-aware design while simultaneously exposing its limits. The inconsistency of broader quality improvements further cautions against treating smell reduction as a proxy for overall quality. In this light, REFINE is best understood as a candidate-generation system rather than an autonomous refactoring tool. Its structured reporting and preservation flags are designed to feed a human-in-the-loop workflow in which compilation, testing, and dependency analysis act as gates before adoption. The taxonomy candidates listed for this work—Architecture, Cybersecurity, Network, and Cryptography—are not directly addressed by the paper's Java refactoring focus, but the preservation-risk findings are relevant to any domain where automated code transformation must not alter security- or behaviour-critical structures.
Who should read this
Opening member content…