Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

SkillSecurer: Detecting and Patching Prompt-Injection Vulnerabilities in AI Agent Skills

A fully agentic red-blue framework achieving 100% injection detection and uncovering latent vulnerabilities in over 17% of popular agent skills
Donato Mecca; Alberto Verna; Youness Bouchari; Nikhil Jha; Marco Melliaยท 2026ยท DOI 10.48550/arXiv.2609.14079

The core problem

Agent skills extend AI agents with reusable instructions, scripts, and configuration, enabling modular and composable behaviour. However, this extensibility introduces a new attack surface: malicious or compromised skills can influence an agent's decisions and actions through prompt injection. Existing security scanners often flag skills at a coarse level, lacking localisation and actionable remediation. This paper presents SkillSecurer, a fully agentic framework designed to generate, detect, localise, and remediate security risks in agent skills. The framework addresses the need for context-aware analysis that goes beyond simple flagging, providing grounded evidence and patches. The authors evaluate SkillSecurer by selecting the best backend LLM, comparing it with competitors, and manually cross-validating each evaluation stage. The results demonstrate that SkillSecurer is the only scanner to achieve a 100% injection detection rate with its best-performing backend. Furthermore, an analysis of popular skills from skills.sh reveals latent vulnerabilities in more than 17% of the skills examined, and testing some of those skills triggered actual incidents, underscoring the risks of ru

Innovation

The evaluation of SkillSecurer yields several key findings. With its best-performing backend LLM, SkillSecurer achieves a 100% injection detection rate, outperforming all competitors. The authors compare SkillSecurer with other scanners, demonstrating its superior detection and localisation capabilities. Manual cross-validation confirms the reliability of the evaluation stages. In a real-world analysis of popular skills from skills.sh, SkillSecurer uncovers latent vulnerabilities in more than 17% of the skills examined. Testing some of these vulnerable skills triggers actual incidents, providing concrete evidence of the risks associated with running unverified skills. These results highlight the prevalence of prompt-injection vulnerabilities in the agent skill ecosystem and the effectiveness of SkillSecurer in identifying them. The framework's ability to localise injections and propose patches enables actionable remediation, distinguishing it from scanners that only flag skills at a high level.
Agent skills extend AI agents with reusable instructions, scripts, and configuration, enabling modular and composable behaviour. However, this extensibility introduces a new attack surface: malicious or compromised skills can influence an agent's decisions and actions through prompt injection. Existing security scanners often flag skills at a coarse level, lacking localisation and actionable remediation. This paper presents SkillSecurer, a fully agentic framework designed to generate, detect, localise, and remediate security risks in agent skills. The framework addresses the need for context-aware analysis that goes beyond simple flagging, providing grounded evidence and patches. The authors evaluate SkillSecurer by selecting the best backend LLM, comparing it with competitors, and manually cross-validating each evaluation stage. The results demonstrate that SkillSecurer is the only scanner to achieve a 100% injection detection rate with its best-performing backend. Furthermore, an analysis of popular skills from skills.sh reveals latent vulnerabilities in more than 17% of the skills examined, and testing some of those skills triggered actual incidents, underscoring the risks of running unverified skills.
SkillSecurer employs a dual-agent architecture: a red agent and a blue agent, with an optional verifier for controlled instances. The red agent generates context-compatible injections across nine threat types while recording the exact modification made to the skill. This recording enables injection-level evaluation. The blue agent analyses complete skill packages, produces grounded evidence, and proposes patches to remediate detected vulnerabilities. For controlled instances, a verifier compares the blue agent's findings and patches with the red agent's recorded injection, allowing precise assessment of detection and localisation accuracy. The framework is fully agentic, leveraging LLMs for both generation and analysis. The authors select the best backend LLM through evaluation and compare SkillSecurer with competing scanners. Manual cross-validation is performed at each evaluation stage to ensure reliability. The methodology emphasises context-aware analysis, enabling reliable injection localisation and actionable remediation beyond skill-level flagging alone.

Why it matters

The findings underscore the importance of context-aware LLM analysis for securing agent skills. SkillSecurer's dual-agent approach, combined with a verifier for controlled instances, provides a robust mechanism for generating, detecting, and remediating prompt-injection vulnerabilities. The 100% detection rate achieved with the best backend LLM demonstrates the potential of agentic frameworks in this domain. The discovery of latent vulnerabilities in over 17% of popular skills indicates a widespread security gap in the emerging agent skill ecosystem. The actual incidents triggered by testing vulnerable skills emphasise the real-world consequences of running unverified skills. The authors argue that reliable injection localisation and actionable remediation are essential beyond simple skill-level flagging. SkillSecurer's ability to produce grounded evidence and patches facilitates trust and safety in agent skill deployment. Future work may involve expanding the threat taxonomy, improving patch generation, and integrating SkillSecurer into skill marketplaces for continuous security assessment. The framework's design aligns with the need for automated, scalable security solutions in AI agent ecosystems.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ