Ilmu Komputer & AI editorial
Open AccessOA2026
ActGov: Governing LLM Agent Actions via Policy-Constrained Validation
A runtime enforcement framework that validates each LLM-proposed tool action before it causes external effects, using SMT-verified policy sets and per-action authorization boundaries.
Kaiyuan Zhang; Yuke Peng; Ke Jiang; Yinqian Zhangยท 2026ยท DOI 10.48550/arXiv.2609.24446
The core problem
Large language model (LLM) agents increasingly execute long-horizon workflows through external tools, allowing untrusted outputs to influence subsequent actions and exceed user authorization. Existing defenses isolate injected content or constrain execution with predefined plans and static policies, but these approaches are brittle under dynamic workflows and scale poorly across extensible tool ecosystems. The core problem is that an agent may propose a tool action that is semantically plausible but violates the user's authorization boundary, especially when untrusted tool outputs inject malicious instructions. ActGov addresses this by validating each LLM-proposed tool action before it causes external effects, rather than relying on the underlying LLM to correctly identify malicious instructions. The framework is built on a unified semantic model of authorization, actions, runtime context, and security constraints, enabling fine-grained enforcement over dynamic agent executions.
Innovation
ActGov is evaluated on the AgentDojo and AgentDyn benchmarks across multiple models and attack configurations. The results show that ActGov consistently reduces the success rate of indirect prompt-injection attacks while preserving task utility, significantly outperforming existing defenses. Specifically, ActGov achieves lower attack success rates than baseline defenses under various attack configurations, without sacrificing the agent's ability to complete benign tasks. The evaluation spans multiple models, demonstrating that the enforcement mechanism is model-agnostic and does not rely on the LLM's internal ability to detect malicious instructions. These results demonstrate that ActGov can enforce fine-grained authorization over dynamic agent executions, providing a robust defense against indirect prompt injection in long-horizon workflows.
Large language model (LLM) agents increasingly execute long-horizon workflows through external tools, allowing untrusted outputs to influence subsequent actions and exceed user authorization. Existing defenses isolate injected content or constrain execution with predefined plans and static policies, but these approaches are brittle under dynamic workflows and scale poorly across extensible tool ecosystems. The core problem is that an agent may propose a tool action that is semantically plausible but violates the user's authorization boundary, especially when untrusted tool outputs inject malicious instructions. ActGov addresses this by validating each LLM-proposed tool action before it causes external effects, rather than relying on the underlying LLM to correctly identify malicious instructions. The framework is built on a unified semantic model of authorization, actions, runtime context, and security constraints, enabling fine-grained enforcement over dynamic agent executions.
ActGov consists of two main components: ActGov-Policy and ActGov-Runtime. ActGov-Policy iteratively constructs a policy set from tool specifications, benign tasks, and observed failure traces. Each policy update is verified through SMT-based counterexample checking, ensuring that the policy set remains sound with respect to the authorization model. Formally, let
be the set of actions,
the runtime context, and
the policy set. A tool call
is permitted only if it satisfies all applicable policies and remains within the task-scoped authorization boundary
. The runtime enforcement condition can be expressed as:
Why it matters
The key insight of ActGov is that runtime enforcement at the action level is more robust than static planning or content isolation. By abstracting each tool call into finite policy records and checking them against a verified policy set, ActGov ensures that even if an untrusted output influences the agent's proposal, the action will be blocked if it violates the authorization boundary. The SMT-based counterexample checking provides formal guarantees that the policy set is consistent with the intended authorization model, reducing the risk of policy gaps. However, the approach relies on the completeness of the semantic model and the quality of failure traces; incomplete specifications could lead to missed policies. Future work could extend the framework to handle more complex authorization models and to automatically synthesize policies from natural language specifications. Overall, ActGov represents a significant step toward secure and trustworthy LLM agent deployments in dynamic, tool-rich environments.
Who should read this
CS practitioners and researchers
Opening member contentโฆ