Computer Science editorial
Large Language Models for Requirements Engineering: A Cross-Task Empirical Evaluation
The core problem
Requirements-related information is scattered across heterogeneous artefacts such as user feedback, developer discussions, and software repositories, making the extraction of actionable requirements knowledge labour-intensive and hard to scale. Large Language Models (LLMs) can support many Requirements Engineering (RE) activities, from classification and traceability identification to specification and explanation generation. However, existing evidence is fragmented across tasks, artefact types, and evaluation settings, and studies rarely offer cross-task evaluations or replication packages.
To address this gap, the authors present two complementary empirical studies evaluating LLMs across five RE-related activities. The first is a controlled experiment on five lightweight open-source LLMs for feedback-driven requirements classification and specification generation. The second is an exploratory industrial case study on two frontier LLMs for traceability link identification and traceability explanation generation using real project artefacts. Classification and traceability identification were assessed with quantitative metrics, while generation tasks were evaluated through human e
Innovation
The abstract reports that LLM performance is strongly task-dependent, ranging from moderate to high. No single model consistently outperformed the others across the evaluated activities. This indicates that effective adoption depends on selecting models and prompting strategies per task.
Quantitative results for classification and traceability identification are not detailed in the abstract beyond the qualitative summary. Similarly, human evaluation outcomes for specification generation and traceability explanation generation are summarised as part of the overall task-dependent performance pattern. The key empirical finding is the absence of a universally dominant model, which underscores the need for task-specific model selection.
Because the source abstract does not provide per-task scores, model names, or dataset statistics, this digest cannot report specific numerical results without inventing them. The replication materials mentioned as a contribution are intended to allow others to reproduce and extend the quantitative and human-evaluation findings.
Why it matters
The cross-task evaluation reveals that LLM capabilities for RE are not monolithic. Performance varies substantially by activity, and the lack of a consistently superior model implies that practitioners should not assume that a single frontier or lightweight model will suffice for all RE tasks. Instead, adoption strategies should be task-aware, pairing models and prompting strategies to the specific RE activity.
The two-study design strengthens the contribution by combining controlled experimentation on lightweight open-source models with an exploratory industrial case study on frontier models using real project artefacts. This dual approach balances internal validity with ecological relevance. The inclusion of replication materials addresses a known weakness in the literature: fragmented evidence and rare replication packages.
Limitations include the exploratory nature of the industrial case study and the reliance on human evaluation for generation tasks, which can introduce subjectivity. The abstract does not report effect sizes, confidence intervals, or inter-rater agreement, so the robustness of the generation-task findings cannot be assessed from the available information. Future work could expand the taxonomy of RE activities, include more models, and standardise evaluation protocols.
From a practical readiness perspective, the results suggest that current LLMs are moderately to highly capable for some RE activities but require careful model selection and prompting. Organisations should pilot task-specific configurations rather than deploying a single model across the RE lifecycle.
Who should read this
Opening member contentโฆ