Computer Science editorial
Open AccessOA2026
Automating Attack Graph Construction for Agentic Pentesting: Towards Neuro-Symbolic Vulnerability Hunting
A semi-automated pipeline translating scanner evidence into MulVAL predicates and LLM-assisted Datalog rules for auditable attack path reasoning
Oliver Stevanovic; Jasmin Wachterยท 2026ยท DOI 10.48550/arXiv.2609.15523
The core problem
Logic attack graphs grounded in scanner output provide explicit and auditable attack path reasoning that LLM-based agents lack. Integrating symbolic frameworks such as MulVAL to contemporary security workflows or agentic pipelines, however, requires translating scanner evidence to initial facts, and creating domain-specific rules. The authors address this interoperability problem by presenting a semi-automated pipeline and demonstrating its feasibility in a web-security case study. The core motivation is to combine the strengths of neuro-symbolic AI: the pattern recognition and flexibility of LLMs with the rigorous, explainable inference of symbolic systems like MulVAL/XSB. This work aims to bridge the gap between raw scanner outputs and formal attack graph construction, enabling more transparent and verifiable vulnerability hunting in agentic pentesting scenarios.
Innovation
Every task produced at least one goal-reaching graph. The pipeline achieved a mean ground-truth vulnerability coverage of 53.7%, with 51.9% of tasks achieving full coverage. The mean noise-path rate was 83.9%, indicating a high proportion of paths that do not correspond to actual vulnerabilities. The median end-to-end time was 24.9 seconds, with MulVAL reasoning taking 2.7 seconds. These results demonstrate that the pipeline is feasible and runtime-practical for agentic workflows. However, predicate coverage, rule coverage, and path precision remain limiting factors. The high noise-path rate suggests that while the system can generate attack graphs, many paths are spurious, which could hinder downstream agent performance if not filtered.
Logic attack graphs grounded in scanner output provide explicit and auditable attack path reasoning that LLM-based agents lack. Integrating symbolic frameworks such as MulVAL to contemporary security workflows or agentic pipelines, however, requires translating scanner evidence to initial facts, and creating domain-specific rules. The authors address this interoperability problem by presenting a semi-automated pipeline and demonstrating its feasibility in a web-security case study. The core motivation is to combine the strengths of neuro-symbolic AI: the pattern recognition and flexibility of LLMs with the rigorous, explainable inference of symbolic systems like MulVAL/XSB. This work aims to bridge the gap between raw scanner outputs and formal attack graph construction, enabling more transparent and verifiable vulnerability hunting in agentic pentesting scenarios.
The proposed pipeline consists of three main stages: (1) parsing findings from Trivy, Semgrep, and Nmap into MulVAL predicates; (2) using an LLM-assisted process to construct domain-specific Datalog rules that link scanner-detectable evidence to attack techniques; and (3) performing symbolic inference with MulVAL/XSB to generate structured attack paths. The pipeline is semi-automated, meaning human oversight may be involved in rule validation. The LLM assists in translating natural language descriptions of vulnerabilities into formal Datalog rules, reducing manual effort. The architecture can be represented as follows:
Why it matters
The study highlights the potential of neuro-symbolic approaches for automated attack graph construction, but also reveals significant challenges. The 53.7% mean coverage indicates that roughly half of the ground-truth vulnerabilities are captured, leaving room for improvement in predicate and rule coverage. The 83.9% noise-path rate is particularly concerning, as it implies that the majority of generated paths are not useful, which could overwhelm analysts or agents. The authors suggest next steps including semantic rule validation and agent-level comparison for graph-guided pentesting. The integration of LLMs for rule generation is promising but requires validation to ensure correctness and completeness. The runtime performance (median 24.9 s) is acceptable for many agentic workflows, but further optimization may be needed for real-time applications. Overall, the work provides a foundation for combining symbolic reasoning with LLM-based agents, but more research is needed to improve precision and coverage.
Who should read this
CS practitioners and researchers
Opening member contentโฆ