Computer Science editorial
Open AccessOA2026
GhostSplat: Input-Triggered Backdoors for Multi-View-Consistent 3D Content Manipulation in Feed-Forward Gaussian Splatting
A supply-chain attack that installs persistent, multi-view-consistent payloads into shared feed-forward 3DGS generator weights
Yudong Gao; Zongjian Ding; Linghan Chen; Yajing Chen; Yu Xinglin; Jiale Liu; Shan Huang; Mingjun Chengยท 2026ยท DOI 10.48550/arXiv.2608.29184
The core problem
Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a 3D scene from sparse images in a single forward pass, enabling fast, generalizable novel-view synthesis. Because these systems rely on shared pretrained generator weights, they introduce a supply-chain attack surface: a compromised checkpoint can affect every downstream scene rendered by the model. Prior backdoor work on Neural Radiance Fields (NeRF) and per-scene 3DGS modifies individual scenes and activates at selected viewpoints, but does not install persistent behavior in shared generator weights. GhostSplat addresses this gap by introducing an input-triggered backdoor for feed-forward 3DGS. A low-amplitude pattern added to the input images causes the poisoned generator to render an attacker-chosen payload on unseen victim scenes. The payload is anchored to a 3D point and reprojected into each target view, making it multi-view consistent. The authors observe that exact projection onto the generator's representation-specific consistency set leaves a realized payload unchanged, because the output already belongs to that set. This insight frames both the attack's stealth and the limitations of naive defenses.
Innovation
GhostSplat succeeds across three feed-forward 3DGS architectures (MVSplat, pixelSplat, DepthSplat) and two datasets (RealEstate10K, ACID). In the strongest evaluated injection and deletion settings, the attack reaches 96% and 100% attack success rate (ASR), respectively, with zero observed false positives. The payload remains multi-view consistent because it is anchored to a 3D point and reprojected into each target view. Robustness evaluations show the backdoor survives JPEG compression, blur, and resampling, indicating that the trigger persists under common image transformations. The zero false-positive rate means clean inputs do not trigger the payload, preserving normal reconstruction behavior. These results demonstrate that a low-amplitude input pattern can reliably install persistent behavior in shared generator weights, affecting unseen victim scenes rather than only the scenes seen during poisoning. The consistency of results across architectures and datasets suggests the vulnerability is not specific to a single model or data distribution.
Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a 3D scene from sparse images in a single forward pass, enabling fast, generalizable novel-view synthesis. Because these systems rely on shared pretrained generator weights, they introduce a supply-chain attack surface: a compromised checkpoint can affect every downstream scene rendered by the model. Prior backdoor work on Neural Radiance Fields (NeRF) and per-scene 3DGS modifies individual scenes and activates at selected viewpoints, but does not install persistent behavior in shared generator weights. GhostSplat addresses this gap by introducing an input-triggered backdoor for feed-forward 3DGS. A low-amplitude pattern added to the input images causes the poisoned generator to render an attacker-chosen payload on unseen victim scenes. The payload is anchored to a 3D point and reprojected into each target view, making it multi-view consistent. The authors observe that exact projection onto the generator's representation-specific consistency set leaves a realized payload unchanged, because the output already belongs to that set. This insight frames both the attack's stealth and the limitations of naive defenses.
GhostSplat trains a poisoned feed-forward 3DGS generator by pairing a low-amplitude input trigger with a 3D-anchored payload. The trigger is a subtle pattern added to input images; when present, the generator renders the attacker-chosen payload on unseen victim scenes. To ensure the payload appears consistently across viewpoints, the attacker anchors it to a 3D point and reprojects it into each target view. Formally, let the generator be and the trigger be . For a clean input set , the poisoned output is
, where
is the reconstructed scene and is the payload anchored at 3D point . Reprojection into view yields , ensuring multi-view consistency. The training objective combines a reconstruction loss on clean inputs with a payload injection loss on triggered inputs, while constraining the trigger amplitude to remain low. The authors evaluate the framework across three architectures (MVSplat, pixelSplat, DepthSplat) and two datasets (RealEstate10K, ACID). The attack is assessed under injection and deletion settings, with robustness tests against JPEG compression, blur, and resampling. The key theoretical observation is that exact projection onto the generator's representation-specific consistency set leaves a realized payload unchanged, because the output already belongs to that set.
Why it matters
The central security implication is that defenses relying solely on exact projection onto the generator's representation-specific consistency set are insufficient. Because a realized payload already belongs to that set, projecting it back onto the set leaves it unchanged, so such defenses cannot distinguish the poisoned output from a legitimate one. Effective mitigation therefore requires information or intervention beyond same-set consistency projection. This finding reframes the threat model for feed-forward 3DGS: shared pretrained weights are a supply-chain vector, and input-triggered backdoors can install persistent, multi-view-consistent payloads that generalize to unseen scenes. The attack's robustness to JPEG, blur, and resampling further complicates detection. Future defenses may need to inspect training provenance, perturb inputs beyond the trigger distribution, or enforce consistency constraints that are not satisfied by the payload by construction. The work also highlights a broader tension in 3D generative systems: the same consistency properties that make reconstructions coherent can be exploited to make malicious payloads coherent as well.
Who should read this
CS practitioners and researchers
Opening member contentโฆ