Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

Loopjacking: Hijacking Human-in-the-Loop Approval

A security analysis of approval-binding failures in agentic AI systems
Adithyan Arun Kumarยท 2026ยท DOI 10.48550/arXiv.2609.21081

The core problem

Human-in-the-loop (HITL) approval is widely regarded as the final security barrier before an autonomous agent performs a consequential action. However, this barrier is only effective if the operation presented for human review is identical to the operation that is later authorized and executed. The paper identifies a class of vulnerabilities termed **Loopjacking**, where the binding between the approved operation and the executed operation is broken. Formally, let be the operation the human believes they are approving, and be the operation that is actually executed. Loopjacking occurs when the human's approval decision for is used to authorize , with in terms of security-relevant semantics. The authors distinguish two variants: **representation-based attacks**, where is already encoded but omitted or misrepresented at approval time, and **post-approval state-substitution attacks**, where the human sees the correct but mutable workflow state later replaces it with . The paper evaluates a purposive set of released agent products to demonstrate the feasibility of these attacks and to test potential mitigations.

Innovation

The evaluation successfully reproduced Loopjacking attacks in multiple products. Post-approval state substitution was reproduced in seven tested Agno AgentOS releases (up to 3.0.9) and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition (up to 0.14.0). Representation mismatch was reproduced in OpenClaw 2026.2.23, and the attack was rejected in the subsequent release 2026.2.24, indicating a fix. The OpenAI Agents SDK versions 0.22.0 and 0.22.2 served as a negative control: their serialized continuation mechanism preserved exact per-call binding and rejected mutated operations . This demonstrates that secure binding is achievable. The results are summarized in the following table:

| Product | Version(s) | Attack Type | Outcome |
|---------|------------|-------------|---------|
| Agno AgentOS | โ‰ค 3.0.9 | Post-approval substitution | Reproduced |
| LangGraph Agent Server | โ‰ค 0.14.0 | Post-approval substitution | Reproduced |
| OpenClaw | 2026.2.23 | Representation mismatch | Reproduced |
| OpenClaw | 2026.2.24 | Representation mismatch | Rejected (fixed) |
| OpenAI Agents SDK | 0.22.0, 0.22.2 | Both | Rejected (negative control) |

These findings hig

Human-in-the-loop (HITL) approval is widely regarded as the final security barrier before an autonomous agent performs a consequential action. However, this barrier is only effective if the operation presented for human review is identical to the operation that is later authorized and executed. The paper identifies a class of vulnerabilities termed **Loopjacking**, where the binding between the approved operation and the executed operation is broken. Formally, let be the operation the human believes they are approving, and be the operation that is actually executed. Loopjacking occurs when the human's approval decision for is used to authorize , with in terms of security-relevant semantics. The authors distinguish two variants: **representation-based attacks**, where is already encoded but omitted or misrepresented at approval time, and **post-approval state-substitution attacks**, where the human sees the correct but mutable workflow state later replaces it with . The paper evaluates a purposive set of released agent products to demonstrate the feasibility of these attacks and to test potential mitigations.
The authors conducted a purposive evaluation of several released agent products, focusing on their HITL approval mechanisms. They tested for two types of Loopjacking: representation mismatch and post-approval state substitution. The evaluation involved reproducing attacks in controlled environments and observing whether the systems allowed the execution of an operation different from the one approved. The specific products tested include:

Why it matters

The paper's core contribution is the identification and formalization of Loopjacking as a distinct class of approval-binding failures. It separates this from related work on misleading dialogs, session smuggling, action binding, and authorization continuity. The attacks exploit the gap between the human's mental model of the approved operation and the actual operation executed. The authors propose two mitigation strategies: (1) complete canonical approval rendering, which ensures the human sees the exact operation to be executed, and (2) exact use-time comparison, which verifies that the operation at execution time matches the approved one. Alternatively, preventing unauthorized mutation of pending state can also block post-approval substitution. The negative control with OpenAI Agents SDK shows that serialized continuation can preserve binding. The results do not estimate ecosystem prevalence but demonstrate that the vulnerability exists in popular frameworks. The authors emphasize that HITL approval is only meaningful if the binding is secure. The following Mermaid diagram illustrates the Loopjacking attack flow and the mitigation points:

The paper concludes that securing the approval binding is essential for trustworthy agentic systems.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ