Ilmu Komputer & AI editorial
SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing
The core problem
Innovation
The proposed architecture, SpliTEE, splits LLM inference between a TEE (CPU-based) and an untrusted GPU. The TEE holds the model weights and performs sensitive operations, while the GPU handles computationally intensive matrix multiplications. Intermediate activations sent to the GPU are masked using differential privacy. Specifically, the TEE adds calibrated noise to the intermediate representations before outsourcing them. The noise scale is determined by a global sensitivity analysis of the LLM functions, which bounds the maximum change in output due to a single input change. This analysis is crucial for ensuring that the DP guarantee holds. The authors derive an upper bound on the floating-point error introduced by masking and noise cancellation in the TEE as a function of the privacy parameter . The DP mechanism used is likely the Gaussian mechanism, where noise
Why it matters
Who should read this
Opening member contentโฆ