Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing

A hardware-software co-designed framework integrating near-data processing with PCA-guided early exiting and dynamic-float compression for memory-bound ANNS in RAG systems.
Cheng Zou; Shuo Yang; Chen Nie; Yu Zou; Yu He; Chao Jiang; Limin Xiao; Weifeng Zhang; Zhezhi He· 2026· DOI 10.48550/arXiv.2605.21952

The core problem

Large language models (LLMs) have advanced rapidly, and retrieval-augmented generation (RAG) has become the key mechanism for expanding model knowledge and reducing hallucinations. Central to RAG is approximate nearest neighbor search (ANNS), which retrieves database vectors most similar to a given query. However, distance calculation over high-dimensional vectors is inherently memory-bound, causing retrieval performance to be constrained by I/O bandwidth on mainstream platforms such as CPUs and GPUs. Although many prior early exiting (EE) techniques attempt to reduce memory accesses by only computing partial dimensions, the partial distance converges too slowly to the EE threshold, which ultimately limits their performance gains. To address these challenges, the authors propose NASZIP, a hardware-software co-designed framework that integrates near data processing (NDP) with a novel feature-level early exiting guided by statistics-based principal component analysis (PCA). Instead of relying solely on partial distances, NASZIP incorporates estimation and correction parameters to approximate full dimensional distances accurately, enabling earlier exiting without compromising accuracy

Innovation

With these co-optimized techniques, NASZIP delivers speedups of up to / over CPU baseline and state-of-the-art GPU implementation at equal accuracy. Relative to the state-of-the-art NDP ANNS accelerator ANSMET, NASZIP achieves performance improvement. These results demonstrate the effectiveness of combining near-data processing with feature-level early exiting and dynamic-float compression. The speedup over CPU is particularly notable, as CPUs are commonly used in RAG pipelines but suffer from limited memory bandwidth. The speedup over a state-of-the-art GPU implementation shows that NASZIP can outperform even highly optimized GPU-based ANNS, which are typically memory-bound. The improvement over ANSMET, a dedicated NDP accelerator, highlights the benefits of the proposed PCA-guided early exiting and dynamic-float scheme beyond existing NDP approaches. All speedups are achieved at equal accuracy, meaning that the approximation introduced by early exiting and compression does not degrade retrieval quality.
Large language models (LLMs) have advanced rapidly, and retrieval-augmented generation (RAG) has become the key mechanism for expanding model knowledge and reducing hallucinations. Central to RAG is approximate nearest neighbor search (ANNS), which retrieves database vectors most similar to a given query. However, distance calculation over high-dimensional vectors is inherently memory-bound, causing retrieval performance to be constrained by I/O bandwidth on mainstream platforms such as CPUs and GPUs. Although many prior early exiting (EE) techniques attempt to reduce memory accesses by only computing partial dimensions, the partial distance converges too slowly to the EE threshold, which ultimately limits their performance gains. To address these challenges, the authors propose NASZIP, a hardware-software co-designed framework that integrates near data processing (NDP) with a novel feature-level early exiting guided by statistics-based principal component analysis (PCA). Instead of relying solely on partial distances, NASZIP incorporates estimation and correction parameters to approximate full dimensional distances accurately, enabling earlier exiting without compromising accuracy. The work further introduces a bit-level NDP-aware dynamic-float scheme that significantly reduces memory access for vector data. On the hardware side, a data aware neighbor list mapping strategy reduces neighbor retrieval latency and inter-channel communication overhead, complemented by a dedicated cache that exploits data locality and enhances prefetch efficiency.
NASZIP is a hardware-software co-designed framework that integrates near-data processing (NDP) with a novel feature-level early exiting guided by statistics-based principal component analysis (PCA). The core idea is to avoid computing full-dimensional distances for every candidate vector by using PCA to identify the most informative dimensions and to estimate the remaining distance contributions. Instead of relying solely on partial distances, NASZIP incorporates estimation and correction parameters to approximate full dimensional distances accurately, enabling earlier exiting without compromising accuracy. This feature-level early exiting is guided by statistics derived from PCA, which allows the system to determine when a partial distance is sufficiently close to the full distance to safely exit computation.

Why it matters

The key insight of NASZIP is that partial distance converges too slowly to the early exiting threshold in prior EE techniques, limiting their performance gains. By using PCA to guide feature-level early exiting and incorporating estimation and correction parameters, NASZIP approximates full-dimensional distances accurately, enabling earlier exiting without compromising accuracy. This addresses the fundamental bottleneck of memory-bound distance calculation in ANNS. The bit-level NDP-aware dynamic-float scheme further reduces memory access by compressing vector data in a way that is aware of the NDP architecture. The hardware-side optimizations—data-aware neighbor list mapping and a dedicated cache—reduce neighbor retrieval latency and inter-channel communication overhead, and enhance prefetch efficiency. Together, these techniques yield substantial speedups over CPU, GPU, and the state-of-the-art NDP accelerator ANSMET. The results suggest that co-designing software algorithms with NDP hardware is a promising direction for accelerating ANNS in RAG systems. Future work may explore extending the approach to other memory-bound workloads and integrating with emerging memory technologies. The taxonomy candidates for this work include Architecture, Cybersecurity, Network, and Cryptography, though the primary contribution lies in computer architecture and hardware acceleration for machine learning systems.

Who should read this

CS practitioners and researchers

Opening member content…