Ilmu Komputer & AI editorial
When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents
The core problem
Harness design has transformed the development of LLM-based agents by integrating memory, tool use, and runtime control. This architectural shift enables agents to maintain context, invoke external tools, and operate across sessions, but it also introduces security and privacy risks. In particular, malicious instructions from external sources may be written into persistent memory and persist across sessions, creating a durable attack surface that outlives a single interaction.
The authors propose PMPA (Persistent Memory Poisoning Attack), a novel attack against harness-based agents. PMPA embeds malicious instructions into benign external sources and induces the victim agent to write them into persistent memory without directly accessing the agent framework. Once stored, the poisoned memory can be retrieved in later sessions, triggering additional malicious actions and causing privacy leakage. The paper evaluates PMPA on OpenClaw and Claude Code across different backbone LLMs, input modalities, and trigger scenarios, and further assesses a targeted prompt-level defense.
Innovation
Across all settings, PMPA achieves average ISR and C-ASR of 73.7% and 55.5% on OpenClaw, and 66.9% and 81.7% on Claude Code. These results demonstrate that persistent memory poisoning is a practical and effective attack vector across different harness implementations. The attack preserves benign task performance on both systems, meaning the agent continues to function normally for legitimate tasks while the poisoned memory remains dormant until triggered.
The variation in C-ASR between OpenClaw (55.5%) and Claude Code (81.7%) suggests that harness design choices influence the likelihood of cross-session exploitation. The authors also evaluate a targeted prompt-level defense, finding that it can reduce memory injection in many settings but provides limited protection once the persistent memory has been poisoned. This indicates that prevention at the injection stage is more effective than mitigation after poisoning.
Why it matters
The findings highlight a fundamental tension in harness-based agent design: the same persistent memory that enables continuity and personalization also creates a durable attack surface. Because malicious instructions can be injected indirectly through benign external sources, traditional access controls on the agent framework are insufficient. The attack's cross-session nature means that a single successful injection can compromise all future sessions until the poisoned memory is purged.
The limited effectiveness of prompt-level defenses after poisoning underscores the need for defense-in-depth strategies. These may include memory provenance tracking, anomaly detection on memory writes, and periodic memory sanitization. The authors' evaluation across multiple backbone LLMs and modalities suggests that the vulnerability is not tied to a specific model but is inherent to the harness architecture.
Future work could explore architectural defenses that isolate memory writes, cryptographic integrity checks for persistent memory, and user-facing transparency mechanisms that allow inspection of stored memories. The paper's taxonomy aligns with Cybersecurity and Architecture, as the attack exploits the interplay between agent architecture and security boundaries.
Who should read this
Opening member contentโฆ