Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

Privacy Failure in Split-LLM Training: The Returned Gradient Nullifies the Decoys

A systems-security case study showing that zero gradients from decoy rows leak which rows are real, even when forward-channel privacy and quality checks pass
Georgios Politis; Evangelos Pappasยท 2026ยท DOI 10.48550/arXiv.2609.04382

The core problem

Split-LLM training distributes computation across a Trusted Local Node (TLN) and an Untrusted Cloud Node (UCN). The TLN holds the private loss and sends protected activations to the UCN; the UCN returns its output; the TLN then returns the output gradient. To obscure which rows are real, the frame the UCN receives mixes real rows with decoys, and the loss ignores the decoys. The authors, Georgios Politis and Evangelos Pappas, present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested. The central observation is simple and severe: because the loss ignores decoys, their gradients are exactly zero, so the pattern of zeros reveals which rows were real. The paper reports measurements from a protocol fixed in advance, including a leak injected at known strength to prove the instrument can see one, a shuffled-label control to prove it does not report absent leaks, and a threshold set before the runs. The work is framed as a case study rather than a complete security proof, and the authors explicitly note that five classes of attack, including those accumulating observations across training step

Innovation

Across nine seeds, the zeros identified the real rows on every frame, 4,096 of 4,096 per run. An attack on the frame contents recovered about one extra token per hundred over a constant-guess baseline, with gains ranging from +0.65 to +1.50 percentage points. The shuffled controls recovered nothing, confirming that the instrument does not report absent leaks. A second set of runs repeated this on a configuration that keeps model quality within budget, so the finding is not confined to a setting nobody would deploy. On both datasets, every such run passed the forward-channel privacy check and the quality check, yet failed that same check once the returned gradient was included. Clipping and noising each row of the gradient closed the leak for about 0.01 nats of held-out cross-entropy. The authors emphasize that the system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured. The results are striking because the leak is exact and deterministic at the row level: the zero pattern is not a statistical artifact but a direct consequence of the loss ignoring decoys.
Split-LLM training distributes computation across a Trusted Local Node (TLN) and an Untrusted Cloud Node (UCN). The TLN holds the private loss and sends protected activations to the UCN; the UCN returns its output; the TLN then returns the output gradient. To obscure which rows are real, the frame the UCN receives mixes real rows with decoys, and the loss ignores the decoys. The authors, Georgios Politis and Evangelos Pappas, present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested. The central observation is simple and severe: because the loss ignores decoys, their gradients are exactly zero, so the pattern of zeros reveals which rows were real. The paper reports measurements from a protocol fixed in advance, including a leak injected at known strength to prove the instrument can see one, a shuffled-label control to prove it does not report absent leaks, and a threshold set before the runs. The work is framed as a case study rather than a complete security proof, and the authors explicitly note that five classes of attack, including those accumulating observations across training steps, were never measured.

The experimental protocol was fixed in advance to avoid post hoc tuning. It includes three components: (1) a leak injected at known strength to prove the instrument can see one, (2) a shuffled-label control to prove it does not report absent leaks, and (3) a threshold set before the runs. The system under test is a two-node split-LLM training setup. The TLN sends protected activations to the UCN; the UCN returns its output; the TLN, holding the private loss, returns the output gradient. The frame the UCN receives mixes real rows with decoys, and the loss ignores the decoys. Formally, for a frame containing real rows and decoy rows , the loss depends only on , so for each decoy row the returned gradient satisfies

, where is the decoy output. The pattern of zeros therefore identifies . The authors run nine seeds and measure the leak across frames. They also run a second set of runs on a configuration that keeps model quality within budget, so the finding is not confined to a setting nobody would deploy. The attack on the frame contents is evaluated against a constant-guess baseline. The mitigation tested is clipping and noising each row of the gradient. The authors report the cost in held-out cross-entropy. The protocol is summarized in the following flow:

Why it matters

The core mechanism is an information leak through the structure of the returned gradient. Because the loss ignores decoys, their gradients are exactly zero, and the pattern of zeros is a perfect indicator of which rows are real. This is not a subtle statistical signal; it is a deterministic channel that survives the forward-channel privacy check. The authors' protocol shows that the leak is measurable and that the shuffled-label control correctly reports no leak when none is present. The attack recovers about one extra token per hundred over a constant-guess baseline, which is a modest but consistent advantage. The second set of runs shows that the leak persists even when model quality is kept within budget, so the finding is not confined to a degenerate configuration. The mitigation of clipping and noising each row of the gradient closes the leak for about 0.01 nats of held-out cross-entropy, which is a small but nonzero cost. However, the authors caution that the system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured. The broader lesson is that privacy evaluations must include the returned gradient as an observable channel, and that decoy-based protections can fail when the loss ignores decoys. The work is a case study, not a complete security proof, and the authors are explicit about the limits of their measurements. The following diagram summarizes the leak:

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ