Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

Reducing Hallucinations in Large Language Models Through Integrated Self-Verification and Retrieval-Augmented Generation

CoVe-RAG+: A Unified Framework for Trustworthy LLMs in Engineering Design
Ashly Joseph· 2026· DOI 10.48550/arXiv.2609.26229

The core problem

Large Language Models (LLMs) are increasingly deployed for advanced engineering tasks such as Computer-Aided Design (CAD) documentation, standards compliance verification, and knowledge retrieval. However, their tendency to generate hallucinations—convincing but unfounded outputs—undermines trustworthiness in high-stakes engineering applications where precision and regulatory compliance are paramount. This paper addresses the critical need for reliable LLM outputs by introducing CoVe-RAG+, a unified framework that combines Chain-of-Verification (CoVe) with Retrieval-Augmented Generation (RAG). The goal is to mitigate hallucinations by grounding LLM responses in authoritative external sources, including engineering standards, CAD data, and simulation reports, while employing iterative self-verification to validate key claims. The work is positioned within the broader context of enhancing LLM reliability for engineering design processes, where factual accuracy is non-negotiable.

Innovation

The authors evaluate CoVe-RAG+ on three engineering tasks: CAD model documentation, standards compliance verification, and reuse of historical design data. Experimental results demonstrate a 28% improvement in factual accuracy compared to baseline CoVe and RAG methodologies. This improvement is measured using a combination of automated fact-checking against ground-truth sources and human evaluation by domain experts. Additionally, CoVe-RAG+ enhances user confidence by providing explanatory verification reports that detail which claims are supported by which sources, thereby offering full traceability. The system also reduces the incidence of hallucinations, particularly for complex queries requiring multi-hop reasoning across standards and design documents. The paper reports that the verification reports were instrumental in helping engineers quickly assess the reliability of LLM outputs, leading to higher adoption rates in design workflows.
Large Language Models (LLMs) are increasingly deployed for advanced engineering tasks such as Computer-Aided Design (CAD) documentation, standards compliance verification, and knowledge retrieval. However, their tendency to generate hallucinations—convincing but unfounded outputs—undermines trustworthiness in high-stakes engineering applications where precision and regulatory compliance are paramount. This paper addresses the critical need for reliable LLM outputs by introducing CoVe-RAG+, a unified framework that combines Chain-of-Verification (CoVe) with Retrieval-Augmented Generation (RAG). The goal is to mitigate hallucinations by grounding LLM responses in authoritative external sources, including engineering standards, CAD data, and simulation reports, while employing iterative self-verification to validate key claims. The work is positioned within the broader context of enhancing LLM reliability for engineering design processes, where factual accuracy is non-negotiable.
CoVe-RAG+ operates through a two-pronged approach: retrieval augmentation and iterative self-verification. First, given a user query, the system retrieves relevant documents from a curated knowledge base comprising engineering standards (e.g., ISO, ASME), CAD model metadata, and simulation reports. These retrieved passages serve as external evidence. Second, the LLM generates an initial response, which is then subjected to a Chain-of-Verification process. This involves decomposing the response into atomic claims, formulating verification questions for each claim, and re-querying the retrieval system to confirm or refute each claim. The verification results are aggregated to produce a final answer, along with a verification report that highlights supported claims, unsupported claims, and source traceability. The process can be formalized as follows: let be the query, the retrieved documents, and the initial response. For each claim in , a verification question is generated, and a retrieval function returns evidence . The claim is accepted if entails ; otherwise, it is revised or rejected. The final response is constructed from verified claims. This iterative loop continues until all claims are either verified or flagged. The framework is designed to be model-agnostic and scalable, integrating seamlessly with existing LLM pipelines.

Why it matters

The findings indicate that integrating self-verification with retrieval augmentation significantly improves the trustworthiness of LLMs in engineering contexts. The 28% accuracy gain underscores the value of grounding LLM outputs in authoritative sources and systematically validating claims. The provision of verification reports addresses the 'black-box' nature of LLMs by making the reasoning process transparent and auditable, which is crucial for compliance-driven industries. However, the approach has limitations: it depends on the quality and coverage of the retrieval corpus, and the iterative verification process may introduce latency. Future work could explore optimizing retrieval efficiency and extending the framework to other domains such as cybersecurity and network design, where similar hallucination risks exist. Overall, CoVe-RAG+ represents a scalable and reliable solution for deploying LLMs in high-stakes engineering design processes, paving the way for broader adoption in safety-critical applications.

Who should read this

CS practitioners and researchers

Opening member content…