Jadwal Sholat

Memuat jadwal sholat…

Ilmu Komputer & AI editorial

Open AccessOA2026

EvoSherlock: Towards Agentic Lifelong Evolution for Unseen Long-Tailed Security-Critical Events in Videos

A causal-enhanced agentic framework for continual classification and temporal localization of emerging security-critical events from scarce samples
Zixin Fan; Jiahong Lu; Changsheng Zheng; Yu Hong; Jingjing Wang· 2026· DOI 10.48550/arXiv.2609.19201

The core problem

Existing Security-oriented Video Understanding (SVU) systems operate under a **closed-world assumption**: static category sets, abundant labels, and all event types known upfront. Real-world security-critical events violate these assumptions—they follow long-tailed distributions, new types emerge continuously, and critical events may offer only a few samples. This paper formalizes the gap as the **Lifelong Evolving Task for Long-Tailed Security-Critical Events in Videos (L²-SCE)**, requiring VLMs to continually classify and temporally localize newly emerging security-critical events from scarce samples without forgetting previously learned events. L²-SCE reveals two critical challenges: (1) **Intra-Event Scarcity**, where extreme data scarcity weakens both classification and temporal localization for new events, and (2) **Inter-Event Interference**, where cross-event feature entanglement and representation drift strengthen catastrophic forgetting. To address these, the authors propose **EvoSherlock**, a causal-enhanced approach orchestrated end-to-end by an **Agentic Controller** with self-reflective closed-loop control, comprising two core modules: the Intra-Event **Causal Video G

Innovation

Extensive experiments on the constructed L²-SCE benchmark demonstrate that EvoSherlock outperforms several advanced baselines in both classification and temporal localization of emerging security-critical events from scarce samples. The benchmark simulates real-world incremental conditions, where new event types appear sequentially with limited annotations. Key quantitative findings include:

- **Classification accuracy** on newly emerged events improves significantly compared to prior methods, validating the effectiveness of CVG in overcoming intra-event scarcity.
- **Temporal localization** (e.g., mean Average Precision) also shows consistent gains, indicating that generated causal samples enhance temporal grounding.
- **Forgetting mitigation**: EvoSherlock exhibits lower backward transfer degradation, confirming that CDA reduces inter-event interference and catastrophic forgetting.

The agentic controller's self-reflective closed-loop control further stabilizes performance across incremental steps. These results justify the importance of the L²-SCE task and the effectiveness of the proposed causal-enhanced agentic approach.

Existing Security-oriented Video Understanding (SVU) systems operate under a **closed-world assumption**: static category sets, abundant labels, and all event types known upfront. Real-world security-critical events violate these assumptions—they follow long-tailed distributions, new types emerge continuously, and critical events may offer only a few samples. This paper formalizes the gap as the **Lifelong Evolving Task for Long-Tailed Security-Critical Events in Videos (L²-SCE)**, requiring VLMs to continually classify and temporally localize newly emerging security-critical events from scarce samples without forgetting previously learned events. L²-SCE reveals two critical challenges: (1) **Intra-Event Scarcity**, where extreme data scarcity weakens both classification and temporal localization for new events, and (2) **Inter-Event Interference**, where cross-event feature entanglement and representation drift strengthen catastrophic forgetting. To address these, the authors propose **EvoSherlock**, a causal-enhanced approach orchestrated end-to-end by an **Agentic Controller** with self-reflective closed-loop control, comprising two core modules: the Intra-Event **Causal Video Generation (CVG)** module and the Inter-Event **Causal Decoupling and Alignment (CDA)** module. A dedicated L²-SCE dataset simulates real-world incremental conditions, and extensive experiments demonstrate EvoSherlock's advantages over advanced baselines, justifying both the task's importance and the method's effectiveness.
EvoSherlock is orchestrated end-to-end by an **Agentic Controller** with self-reflective closed-loop control. The controller dynamically coordinates two causal modules:

Why it matters

The L²-SCE task exposes fundamental limitations of closed-world SVU systems. By formalizing lifelong evolving long-tailed security-critical events, the paper highlights two intertwined challenges: intra-event scarcity and inter-event interference. EvoSherlock's causal-enhanced agentic design directly addresses these through CVG and CDA, orchestrated by a self-reflective controller. The use of causal generation for scarce events is a promising direction for few-shot video understanding, while causal decoupling and alignment offer a principled way to mitigate catastrophic forgetting. The constructed benchmark provides a realistic testbed for future research. However, the approach may face scalability challenges as the number of event types grows, and the computational cost of the agentic loop could be a concern. Future work could explore more efficient controller policies and extend the framework to multimodal security signals. Overall, EvoSherlock represents a significant step toward lifelong, open-world security video understanding.

Who should read this

CS practitioners and researchers

Opening member content…