Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

ExploreAI: Agentic Exploration Knowledge Bases for Reproducible Observable-Regression Testing of Black-Box VR and 3D Applications

An LLM-driven agentic framework that converts exploratory testing traces into reusable, per-object Exploration Knowledge Bases for reproducible observable-regression checking across VR and 3D application versions.
Jiajie Wang; Kebin Peng; Wei Wang; Xiaoyin Wang; Sen He; Xue Qinยท 2026ยท DOI 10.48550/arXiv.2608.21628

The core problem

Black-box VR and 3D applications present a distinctive regression-testing problem: observable failures are not merely a function of code paths but of embodied interaction. Whether a defect becomes visible depends on where a tester moves, which objects enter the field of view, and which views are captured. Manual exploratory testing can surface such failures, but the evidence it produces is time-consuming to reproduce because the exploration path is rarely recorded in a structured, replayable form. Systematic sweeps are reproducible by construction, yet they lack semantic guidance and spend exploration budget on low-value viewpoints that are unlikely to expose meaningful regressions.

The authors observe that a large language model (LLM) can make the same high-level decisions a human tester makes during exploration: interpreting a task, choosing which objects to inspect, grouping related objects, recording what it saw, and deciding when missing evidence should trigger another attempt. Based on this observation, ExploreAI is presented as an LLM-driven agentic framework that offloads repeated perception, navigation, multi-view capture execution, and logging to specialized modules whil

Innovation

Across six indoor and outdoor scenes in Unity, AI2-THOR, and BeamNG, ExploreAI constructs high-completeness EKBs under both complete and target exploration. The completeness result indicates that the agentic loop reliably records the per-object evidence needed for later regression checking, rather than producing partial or inconsistent traces.

The LLM-module ablation shows where semantic planning, capture policy, evidence recording, and self-verification contribute. This ablation design isolates the effect of each LLM-driven decision component, clarifying which parts of the agentic framework are responsible for the observed EKB quality.

Reproduction pilots further show that EKB-guided traces help both humans and LLM-based reproducers reproduce exact object-view evidence more effectively than conditions without EKB context. This is the key practical result: the EKB is not merely a log but a reusable artifact that improves reproducibility of observable evidence across different reproducer types.

Black-box VR and 3D applications present a distinctive regression-testing problem: observable failures are not merely a function of code paths but of embodied interaction. Whether a defect becomes visible depends on where a tester moves, which objects enter the field of view, and which views are captured. Manual exploratory testing can surface such failures, but the evidence it produces is time-consuming to reproduce because the exploration path is rarely recorded in a structured, replayable form. Systematic sweeps are reproducible by construction, yet they lack semantic guidance and spend exploration budget on low-value viewpoints that are unlikely to expose meaningful regressions.
The authors observe that a large language model (LLM) can make the same high-level decisions a human tester makes during exploration: interpreting a task, choosing which objects to inspect, grouping related objects, recording what it saw, and deciding when missing evidence should trigger another attempt. Based on this observation, ExploreAI is presented as an LLM-driven agentic framework that offloads repeated perception, navigation, multi-view capture execution, and logging to specialized modules while reserving the LLM for planning, evidence recording, capture-policy decisions, and verification decisions. The central artifact is the Exploration Knowledge Base (EKB): a structured, per-object record of one exploration run. For each object the agent finds, the EKB stores the scan evidence that exposed it, the selected target, the navigation path, the multi-view capture, and the self-verification result. The EKB is positioned as a reusable testing artifact that supports reproducible observable-regression checking across versions of a VR or 3D application.

Why it matters

The results support the paper's central observation that an LLM can supply the high-level semantic decisions of a human exploratory tester while specialized modules handle repeated perception, navigation, capture, and logging. The high-completeness EKBs under both complete and target exploration suggest that semantic guidance does not come at the cost of coverage, and that the agent can adapt its exploration mode without degrading the structured record it produces.

The ablation is important because it decomposes the framework's contribution: semantic planning, capture policy, evidence recording, and self-verification are each testable components rather than an undifferentiated agent. This makes the design more interpretable and provides a basis for future work to strengthen specific modules.

The reproduction pilots address the core motivation of the paper: manual exploratory evidence is time-consuming to reproduce, while systematic sweeps are reproducible but semantically blind. EKB-guided traces improve reproduction of exact object-view evidence for both humans and LLM-based reproducers, indicating that the structured per-object record captures the information needed to replay observable failures. The EKB thus functions as a bridge between the semantic richness of exploratory testing and the reproducibility of systematic testing.

Limitations follow from the evaluation scope: six scenes across Unity, AI2-THOR, and BeamNG, with reproduction pilots rather than a large-scale industrial deployment. The taxonomy candidates for this work span Architecture, Cybersecurity, Network, and Cryptography, but the paper's direct contribution is to black-box VR and 3D observable-regression testing. Future work could extend EKB-guided reproduction to broader application families and longer version histories.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ