Computer Science editorial
Open AccessOA2026
Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding
CoVeil: A defense mechanism that reduces private-context leakage by up to 87.2% while preserving collaborative decoding quality
Kejia Zhang; Tianyuan Zou; Zixuan GU; Yang Liuยท 2026ยท DOI 10.48550/arXiv.2608.29111
The core problem
Applications such as personalized assistance and proprietary document analysis require large language models (LLMs) to generate outputs from private data. Yet powerful LLMs typically cannot be deployed on the resource-constrained devices where private data resides, and uploading private data to cloud-hosted LLMs exposes sensitive information. Recent work addresses this tension with a cloud-edge collaborative decoding paradigm, where private data are kept on the edge with a small language model (SLM) producing next-token distributions, which are fused with predictions from a cloud LLM operating solely on public data. This paper systematically analyzes the privacy risks of such a paradigm with a novel evaluation framework using constructed QA datasets, showing that such collaboration can expose substantial private-context information. To address this leakage, the authors propose CoVeil, a defense mechanism that dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving collaborative quality. The work sits at the intersection of Architecture, Cybersecurity, Network, and Cryptography, and targets the practical deployment of privacy-preserving LL
Innovation
Extensive evaluations demonstrate that CoVeil consistently improves the privacy-utility trade-off over existing baselines. The mechanism reduces data leakage by up to 87.2%, with minimal accuracy loss. The authors report that the defense preserves collaborative quality while suppressing private-context information, as measured by their constructed QA datasets. The evaluation framework reveals that the unmitigated cloud-edge collaborative decoding paradigm can expose substantial private-context information, underscoring the need for decoding-time defenses. CoVeil's dynamic optimization of transmitted signals achieves a favorable balance between privacy and utility, outperforming prior baselines across the tested scenarios. The results highlight that privacy leakage in collaborative decoding is a practical concern and that targeted signal optimization can effectively mitigate it without sacrificing the benefits of cloud LLM capabilities.
Applications such as personalized assistance and proprietary document analysis require large language models (LLMs) to generate outputs from private data. Yet powerful LLMs typically cannot be deployed on the resource-constrained devices where private data resides, and uploading private data to cloud-hosted LLMs exposes sensitive information. Recent work addresses this tension with a cloud-edge collaborative decoding paradigm, where private data are kept on the edge with a small language model (SLM) producing next-token distributions, which are fused with predictions from a cloud LLM operating solely on public data. This paper systematically analyzes the privacy risks of such a paradigm with a novel evaluation framework using constructed QA datasets, showing that such collaboration can expose substantial private-context information. To address this leakage, the authors propose CoVeil, a defense mechanism that dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving collaborative quality. The work sits at the intersection of Architecture, Cybersecurity, Network, and Cryptography, and targets the practical deployment of privacy-preserving LLM inference.
The authors introduce a novel evaluation framework built on constructed QA datasets to audit privacy leakage in cloud-edge collaborative decoding. In this paradigm, the edge SLM processes private data and produces next-token distributions, while the cloud LLM operates solely on public data; the two distributions are fused to generate the final output. The framework measures how much private-context information is exposed through the transmitted signals.
Why it matters
The paper's analysis centers on the tension between leveraging powerful cloud LLMs and protecting private data on resource-constrained edge devices. By auditing the collaborative decoding paradigm, the authors show that transmitting next-token distributions from an edge SLM can leak private-context information even when raw data never leaves the device. CoVeil addresses this by dynamically optimizing the transmitted signals at decoding time, effectively acting as a privacy filter on the communication channel between edge and cloud. The reported reduction in leakage by up to 87.2% with minimal accuracy loss suggests that the privacy-utility trade-off can be substantially improved over existing baselines. The work contributes both an evaluation framework for auditing leakage and a practical defense mechanism, bridging concerns from Architecture, Cybersecurity, Network, and Cryptography. Future directions may include extending the framework to other collaborative inference settings and further tightening the utility constraint while maintaining strong privacy guarantees.
Who should read this
CS practitioners and researchers
Opening member contentโฆ