Jadwal Sholat

Memuat jadwal sholat…

Ilmu Komputer & AI editorial

Open AccessOA2026

Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

A digest of Shim et al. (2026): why in-context privacy enforcement fails, three new attacks that exploit it, and FLOWSEAL's tool-level information-flow-control defense.
Minsun Shim; Ramisha Raida Karim; Ruthwik Jakkula; Kaiwen Zhou; Xin Liu; Xin Eric Wang; Zhou Li· 2026· DOI 10.48550/arXiv.2609.14003

The core problem

Personal AI agents built on large language models (LLMs) are increasingly granted access to a user's private data and communications to deliver personalized assistance. This access creates a persistent privacy risk: for any given piece of sensitive information, the agent must decide whether disclosure to a particular party is appropriate. Existing defenses attempt to make the backend LLM itself more privacy-preserving through stronger system prompts, fine-tuning, or explicit consent-checking procedures.

The paper identifies a structural flaw in this design philosophy. Whenever enforcement is a judgment the LLM makes over the same conversational context that an adversary controls, the enforcement mechanism and the attack surface coincide. In other words, the very context window used to reason about privacy is also the channel through which an attacker can manipulate that reasoning. The authors argue this coincidence is not a tuning problem but an architectural one, motivating a defense that relocates enforcement outside the model's context entirely.

Formally, let denote the conversational context, the backend LLM, and the privacy policy. In-context defenses compute a

Innovation

FLOWSEAL reduces leak rates to near zero while preserving task utility, regardless of the underlying LLM backend. A headline result is the reduction from 52.2% to 0.5% leak rate against the Collaborative Workspace Lure attack.

The evaluation design is notable for its breadth: three benchmarks, five prompt-based baselines, and eight attacks. Crucially, the evaluation includes a real agent executing live tool calls through MCP, which moves the assessment beyond synthetic prompt-level testing toward realistic agent behavior.

The claim that results hold "regardless of the underlying LLM backend" is significant because it supports the paper's architectural thesis: if enforcement is external to the model, then swapping the model should not reintroduce the vulnerability. This contrasts with in-context defenses, whose effectiveness is entangled with the specific model's instruction-following and refusal behavior.

The near-zero leak rate (0.5%) against an attack that previously succeeded more than half the time (52.2%) represents roughly a two-orders-of-magnitude reduction in successful exfiltration for that attack. The paper reports that this is achieved without sacrificing task utility

Personal AI agents built on large language models (LLMs) are increasingly granted access to a user's private data and communications to deliver personalized assistance. This access creates a persistent privacy risk: for any given piece of sensitive information, the agent must decide whether disclosure to a particular party is appropriate. Existing defenses attempt to make the backend LLM itself more privacy-preserving through stronger system prompts, fine-tuning, or explicit consent-checking procedures.
The paper identifies a structural flaw in this design philosophy. Whenever enforcement is a judgment the LLM makes over the same conversational context that an adversary controls, the enforcement mechanism and the attack surface coincide. In other words, the very context window used to reason about privacy is also the channel through which an attacker can manipulate that reasoning. The authors argue this coincidence is not a tuning problem but an architectural one, motivating a defense that relocates enforcement outside the model's context entirely.

Why it matters

The paper's core analytical contribution is the identification of a structural coincidence: in-context privacy enforcement and the adversarial attack surface occupy the same conversational context. This framing explains why stronger prompts, more training, and consent-checking procedures offer limited robustness—they all operate on , which the adversary can influence.

FLOWSEAL's design follows from this analysis. By moving enforcement to a tool-level interceptor outside the LLM's context, the defense decouples the enforcement mechanism from the attack surface. Data provenance provides the factual basis for decisions, and the information-flow-control lattice provides a principled, checkable policy. Controlled declassification preserves legitimate disclosure when justified by consent or context.

The taxonomy of the work spans several areas: **Architecture** (tool-level interceptor placement, MCP integration), **Cybersecurity** (agent privacy leakage, adversarial interaction), **Network** (channel decoupling across independent channels), and **Cryptography** (information-flow-control lattices, declassification, provenance). The lattice formalism connects the work to classical information-flow-control literature while adapting it to LLM agent tool calls.

Limitations and open questions follow naturally. The interceptor must correctly track provenance across complex tool chains, and declassification policies must be specified by users or platforms. The paper's results suggest these are tractable, but the breadth of real-world tool ecosystems remains a challenge. Nonetheless, the architectural lesson is clear: when enforcement and attack surface coincide, relocate enforcement.

Who should read this

CS practitioners and researchers

Opening member content…