Computer Science editorial
Open AccessOA2026
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
A hardware-accelerated lock agent and transaction agent achieving up to 51ร higher throughput on TPC-C
Shien Zhu; Gustavo Alonsoยท 2026ยท DOI 10.48550/arXiv.2605.13398
The core problem
Online Transaction Processing (OLTP) is a classic application with a growing business. CPU-based OLTP has low lock serving efficiency. The main reason is that most locks are cold, and the lock agent must issue frequent memory accesses to retrieve the lock details to determine whether to grant it. This motivates us to propose dedicated hardware-based lock agents with integrated lock tables to remove the DRAM access overhead. In this paper, we propose hardware-accelerated lock management and transaction processing for database systems. First, we propose a low-latency lock agent optimized for both lock acquiring and releasing requests. Second, we design a scalable transaction agent that executes the full transaction lifecycle. We present the architecture, optimizations, and design-space exploration of the proposed lock management and transaction processing system. The experiment results show up to 51X higher transaction throughput over the CPU baseline on the TPC-C benchmark.
Innovation
The experiment results show up to 51X higher transaction throughput over the CPU baseline on the TPC-C benchmark. This significant improvement is attributed to the elimination of DRAM access overhead for lock management and the parallel processing capabilities of the FPGA. The lock agent achieves low latency for both lock acquiring and releasing requests, reducing contention and improving overall system performance. The transaction agent's scalability allows the system to handle increasing workloads by adding more instances. The design-space exploration reveals trade-offs between resource utilization and performance, with optimal configurations depending on the workload characteristics. The 51X speedup demonstrates the potential of hardware acceleration for OLTP workloads.
Online Transaction Processing (OLTP) is a classic application with a growing business. CPU-based OLTP has low lock serving efficiency. The main reason is that most locks are cold, and the lock agent must issue frequent memory accesses to retrieve the lock details to determine whether to grant it. This motivates us to propose dedicated hardware-based lock agents with integrated lock tables to remove the DRAM access overhead. In this paper, we propose hardware-accelerated lock management and transaction processing for database systems. First, we propose a low-latency lock agent optimized for both lock acquiring and releasing requests. Second, we design a scalable transaction agent that executes the full transaction lifecycle. We present the architecture, optimizations, and design-space exploration of the proposed lock management and transaction processing system. The experiment results show up to 51X higher transaction throughput over the CPU baseline on the TPC-C benchmark.
The proposed system consists of two main hardware components: a lock agent and a transaction agent, both implemented on FPGA. The lock agent integrates lock tables directly into on-chip memory to eliminate DRAM access overhead. It is optimized for low latency in both lock acquisition and release operations. The transaction agent is designed to be scalable and executes the full transaction lifecycle, from request parsing to commit. The architecture is explored through design-space exploration, varying parameters such as the number of lock agents, transaction agents, and on-chip memory sizes. The system is evaluated using the TPC-C benchmark, comparing against a CPU baseline. The key optimization is the removal of frequent memory accesses for cold locks by keeping lock details in integrated lock tables. This hardware-software co-design approach enables high throughput and low latency.
Why it matters
The proposed FPGA-accelerated lock management and transaction processing system addresses the inefficiency of CPU-based OLTP, particularly the overhead of frequent memory accesses for cold locks. By integrating lock tables into hardware, the system removes this bottleneck and achieves substantial throughput gains. The scalability of the transaction agent ensures that the system can be adapted to various workloads. The design-space exploration provides insights into the trade-offs involved in hardware acceleration for database systems. Future work could explore integration with other database components and further optimizations. The results suggest that dedicated hardware accelerators can significantly outperform general-purpose CPUs for specific OLTP tasks, paving the way for more efficient database systems.
Who should read this
CS practitioners and researchers
Opening member contentโฆ