Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

A split-inference architecture that uses differential privacy to protect intermediate representations when outsourcing LLM computation to untrusted GPUs, achieving nearly 2x speedup over CPU-only TEE inference.
Shashie Dilhara Batan Arachchige; Robin Carpentier; Hassan Jameel Asghar; Dali Kaafarยท 2026ยท DOI 10.48550/arXiv.2609.15039

The core problem

Large language models (LLMs) are increasingly deployed as remote services, but user prompts often contain sensitive information. Trusted execution environments (TEEs) offer a solution by isolating computations from the service provider, yet current TEEs are CPU-based and significantly slower than GPUs optimized for LLM inference. Tramer and Boneh (2019) proposed Slalom, which splits neural network inference between a TEE and an untrusted GPU, encrypting intermediate inputs sent to the GPU. However, encryption introduces quantization overhead and limits precision. This work extends split-inference to LLMs and replaces encryption with differential privacy (DP) to protect intermediate representations. The authors first demonstrate the necessity of masking by showing a prompt-reconstruction attack that recovers prompts from intermediate representations with nearly 80% accuracy. They then provide a global sensitivity analysis of key LLM functions to bound the required DP noise scale. Unlike encryption, DP avoids quantization, allowing the LLM to remain in the floating-point domain. The architecture is implemented using Intel TDX and evaluated with Llama-3.2-3B and Qwen3-4B.

Innovation

Experimental evaluation shows that SpliTEE achieves nearly twice the speed of fully CPU-based inference inside Intel TDX. Compared to encryption-based Slalom, SpliTEE is 5-15 seconds faster while achieving higher accuracy. The speedup is attributed to avoiding quantization and enabling floating-point operations on the GPU. The accuracy improvement stems from the DP mechanism preserving the numerical precision of the model, whereas encryption-based approaches often require quantization to fixed-point arithmetic. The authors also demonstrate that prompt reconstruction, even with knowledge of the DP mechanism, cannot recover more information than is contained in an unrelated prompt. This provides a strong privacy guarantee: the masked representations reveal no more about the original prompt than a random prompt would. The prompt-reconstruction attack, which achieved nearly 80% accuracy on unmasked representations, fails to recover meaningful information from DP-masked representations. The floating-point error bound derived in the methodology is validated empirically, showing that the error remains within acceptable limits for the tested values.
Large language models (LLMs) are increasingly deployed as remote services, but user prompts often contain sensitive information. Trusted execution environments (TEEs) offer a solution by isolating computations from the service provider, yet current TEEs are CPU-based and significantly slower than GPUs optimized for LLM inference. Tramer and Boneh (2019) proposed Slalom, which splits neural network inference between a TEE and an untrusted GPU, encrypting intermediate inputs sent to the GPU. However, encryption introduces quantization overhead and limits precision. This work extends split-inference to LLMs and replaces encryption with differential privacy (DP) to protect intermediate representations. The authors first demonstrate the necessity of masking by showing a prompt-reconstruction attack that recovers prompts from intermediate representations with nearly 80% accuracy. They then provide a global sensitivity analysis of key LLM functions to bound the required DP noise scale. Unlike encryption, DP avoids quantization, allowing the LLM to remain in the floating-point domain. The architecture is implemented using Intel TDX and evaluated with Llama-3.2-3B and Qwen3-4B.

The proposed architecture, SpliTEE, splits LLM inference between a TEE (CPU-based) and an untrusted GPU. The TEE holds the model weights and performs sensitive operations, while the GPU handles computationally intensive matrix multiplications. Intermediate activations sent to the GPU are masked using differential privacy. Specifically, the TEE adds calibrated noise to the intermediate representations before outsourcing them. The noise scale is determined by a global sensitivity analysis of the LLM functions, which bounds the maximum change in output due to a single input change. This analysis is crucial for ensuring that the DP guarantee holds. The authors derive an upper bound on the floating-point error introduced by masking and noise cancellation in the TEE as a function of the privacy parameter . The DP mechanism used is likely the Gaussian mechanism, where noise

is added, with calibrated to the sensitivity and :
for -DP. The architecture is implemented using Intel TDX, a CPU-based TEE, and evaluated on two LLMs: Llama-3.2-3B and Qwen3-4B. The split execution is compared against fully CPU-based inference inside TDX and encryption-based Slalom.

Why it matters

The key insight of SpliTEE is that differential privacy can effectively replace encryption in split-inference architectures for LLMs, offering a better trade-off between privacy, performance, and accuracy. By avoiding quantization, DP allows the LLM to operate in its native floating-point domain, which is crucial for maintaining accuracy. The global sensitivity analysis provides a principled way to calibrate the noise scale, ensuring that the privacy guarantee is met without excessive noise that would degrade utility. The derived upper bound on floating-point error as a function of offers a theoretical understanding of the privacy-utility trade-off. The implementation on Intel TDX demonstrates practical feasibility, with significant speedups over CPU-only TEE inference. The comparison with Slalom highlights the advantages of DP over encryption: faster execution and higher accuracy. However, the approach has limitations. The DP guarantee is probabilistic and depends on the choice of and . The sensitivity analysis may be conservative, leading to more noise than necessary. Future work could explore tighter sensitivity bounds, adaptive noise calibration, and extension to other TEEs and LLM architectures. The prompt-reconstruction attack results underscore the importance of masking intermediate representations, and the DP mechanism provides a robust defense even against adversaries aware of the mechanism. Overall, SpliTEE represents a significant step towards practical privacy-preserving LLM inference on trusted hardware with GPU outsourcing.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ