Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

A hybrid architecture that pairs graph attention encoders and PPO policies with LLM narrative generation for enterprise-scale SOC automation
Uday Vallabhaneni; Cassie L. Cagwin; David J. Wildยท 2026ยท DOI 10.48550/arXiv.2609.04159

The core problem

Large language model (LLM) agents are increasingly proposed as autonomous Security Operations Center (SOC) analysts, yet two fundamental limitations undermine their reliability at enterprise scale. First, a finite context window cannot hold a multi-thousand-host authentication graph, which is the essential substrate for understanding lateral movement and containment scope. Second, free-form generation offers no guarantee that a recommended containment action is consistent with the topology it operates onโ€”an LLM may suggest isolating a host that does not exist or severing a connection that would disrupt critical services.

Sentinel-RL addresses these limitations by decoupling topological reasoning from semantic reasoning. The architecture introduces a heterogeneous graph attention encoder that summarizes the live authentication subgraph into a fixed-dimensional state vector. A Proximal Policy Optimization (PPO) policy then maps this state to a constrained set of investigative actions. Finally, an LLM agent loop is restricted to consuming the policy's recommendations and producing analyst-readable narratives, gated by a critic to ensure consistency. This division of labor ensures tha

Innovation

The authors report four primary results. First, the two-phase CREATE ingestion pattern loads a 24M-edge authentication subgraph into Neo4j in 14.2 minutes on a single 32-core node, which is approximately 24x faster than the canonical MERGE-based pipeline (which takes about 5.7 hours). This speedup is critical for real-time SOC operations, where the graph must be continuously updated.

Second, the sliding-window alert engine reliably trips a 25-event/10-second threshold in โ‰ค2.5 s across 50 trials. The engine uses a 10-second window and a threshold of 25 events; the mean detection latency is 1.8 s with a standard deviation of 0.4 s. This ensures that potential attacks are flagged promptly.

Third, PPO training over 200 iterations converges to a mean episodic return of . On held-out labeled red-team events, the policy achieves a precision of 0.91 and a recall of 0.87. The precision-recall curve shows that the policy maintains high precision even at high recall levels, indicating robust performance. The confusion matrix for the held-out set is: true positives = 87, false positives = 9, false negatives = 13, true negatives = 891.

Fourth, the integrated containment loop c

Large language model (LLM) agents are increasingly proposed as autonomous Security Operations Center (SOC) analysts, yet two fundamental limitations undermine their reliability at enterprise scale. First, a finite context window cannot hold a multi-thousand-host authentication graph, which is the essential substrate for understanding lateral movement and containment scope. Second, free-form generation offers no guarantee that a recommended containment action is consistent with the topology it operates onโ€”an LLM may suggest isolating a host that does not exist or severing a connection that would disrupt critical services.
Sentinel-RL addresses these limitations by decoupling topological reasoning from semantic reasoning. The architecture introduces a heterogeneous graph attention encoder that summarizes the live authentication subgraph into a fixed-dimensional state vector. A Proximal Policy Optimization (PPO) policy then maps this state to a constrained set of investigative actions. Finally, an LLM agent loop is restricted to consuming the policy's recommendations and producing analyst-readable narratives, gated by a critic to ensure consistency. This division of labor ensures that topological decisions are made by a model that can ingest the full graph, while the LLM focuses on what it does best: generating human-comprehensible explanations.

Why it matters

The Sentinel-RL architecture demonstrates that offloading topological reasoning from LLM agents to specialized graph neural networks and reinforcement learning policies can significantly improve the reliability and speed of SOC operations. By decoupling topological reasoning from semantic reasoning, the system ensures that containment actions are consistent with the underlying authentication graph, while still leveraging the LLM's strength in generating human-readable narratives. The 24x speedup in graph ingestion is a key enabler for real-time operations, and the two-phase CREATE pattern is a reusable engineering solution to the hot-node deadlock problem.

The PPO policy's performance (precision 0.91, recall 0.87) is promising, but the authors note that false positives still incur costs. At an estimated $50 per false positive in analyst time, the system must be tuned to balance precision and recall based on the organization's risk tolerance. The human-approval boundary is a critical design choice: it ensures reversibility and compliance, but also introduces a median 1.2 s delay. The authors argue that this delay is acceptable given the potential consequences of automated containment errors.

The HPC deployment pattern (anchor-node co-location) is portable to other clusters, but may require adaptation for cloud environments. The enterprise-readiness analysis highlights the importance of audit compliance and reversibility guarantees. Future work could explore multi-agent coordination, where multiple Sentinel-RL instances share a global graph, or incorporate threat intelligence feeds to enrich the graph. Overall, Sentinel-RL represents a significant step toward trustworthy agentic SOC architectures that combine the strengths of graph learning, reinforcement learning, and LLMs while mitigating their individual weaknesses.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ