Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

Origin Is All You Need: Provenance-Aware Transformers for Structural Trust-Boundary Separation

A provenance-aware defense that makes application-supplied source labels actionable inside the model, enforcing a structural boundary between authoritative and non-authoritative sources.
Yuxuan Zhang; Jeff Huang; Guofei Guยท 2026ยท DOI 10.48550/arXiv.2609.21088

The core problem

Indirect prompt injection (IPI) remains a central safety and security challenge for large language model (LLM) systems because standard transformers lack an architectural notion of source authority. Retrieved documents, user inputs, and system instructions are all processed through the same undifferentiated attention mechanism, forcing the model to infer from wording alone what should be obeyed and what should be treated as data. This conflation creates a brittle trust boundary that can be exploited by adversarial content embedded in retrieved documents or user inputs. The authors propose Provenance-Aware Transformers, a provenance-aware defense that makes application-supplied source labels actionable inside the model. By exposing provenance as a first-class architectural signal, the work aims to shift LLM safety alignment from brittle pattern matching toward explicit trust separation.

Innovation

Evaluation shows that Provenance-Aware Transformers maintain robust resistance to IPI both in-distribution and out-of-distribution while preserving utility comparable to the base pretrained model. The authors report that the architecture successfully enforces a structural boundary between authoritative and non-authoritative sources, preventing injected instructions from being obeyed. Quantitative results demonstrate that the model achieves high resistance to IPI attacks without significant degradation in performance on standard tasks. The two-stage fine-tuning pipeline effectively teaches the model origin semantics and task behavior under ring constraints, enabling the model to generalize to unseen attack patterns.
Indirect prompt injection (IPI) remains a central safety and security challenge for large language model (LLM) systems because standard transformers lack an architectural notion of source authority. Retrieved documents, user inputs, and system instructions are all processed through the same undifferentiated attention mechanism, forcing the model to infer from wording alone what should be obeyed and what should be treated as data. This conflation creates a brittle trust boundary that can be exploited by adversarial content embedded in retrieved documents or user inputs. The authors propose Provenance-Aware Transformers, a provenance-aware defense that makes application-supplied source labels actionable inside the model. By exposing provenance as a first-class architectural signal, the work aims to shift LLM safety alignment from brittle pattern matching toward explicit trust separation.
The proposed architecture assigns each input token a ring ID encoding its origin. The model is augmented with three key components: origin embeddings, a learnable origin attention bias, and a learnable origin scale that preserves provenance under normalization. Formally, for a token with origin , the attention score between query and key is modified as:

Why it matters

The work demonstrates that exposing provenance as a first-class architectural signal can shift LLM safety alignment from brittle pattern matching toward explicit trust separation. By making source labels actionable inside the model, the architecture provides a more robust defense against IPI compared to approaches that rely solely on input filtering or prompt engineering. The learnable origin attention bias and scale allow the model to dynamically adjust the influence of different sources during generation, effectively creating a trust boundary that is enforced at the architectural level. This approach is complementary to existing safety measures and can be integrated into released pretrained models via fine-tuning. The authors suggest that future work could explore more granular provenance hierarchies and extend the approach to multi-modal inputs. Overall, the paper presents a promising direction for building more secure and trustworthy LLM systems.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ