Ilmu Komputer & AI editorial
PriMobiBench: Characterizing Visual Privacy Leakage in VLM-Driven Mobile GUI Agents
The core problem
Mobile GUI agents increasingly rely on Vision-Language Models (VLMs) to automate smartphone tasks by interpreting screenshot streams. This design, while powerful, introduces serious and underexplored privacy risks: direct leakage of sensitive on-screen information and unintended user profiling. The absence of standardized benchmarks makes it difficult to quantify these risks in realistic mobile agent workflows.
To address this gap, the authors propose **PriMobiBench**, described as the first benchmark for systematically evaluating privacy leakage and visual profiling in screenshot-driven mobile agents. It provides a unified pipeline for data generation, agent trajectory construction, and multi-model evaluation. The work also introduces **MobiLeak**, a dataset of execution traces from 16 apps, covering 25 privacy attributes with 2,960 embedded privacy instances.
The central research questions are: (1) Can VLMs directly extract sensitive information from screenshots? (2) Can they infer user profiles from aggregated visual evidence? (3) Can a practical mitigation reduce profiling without substantially degrading task performance? The results reveal substantial risks and offer a mitig
Innovation
The evaluation reveals substantial privacy risks in VLM-driven mobile GUI agents.
**Direct leakage.** VLMs can directly extract sensitive information with up to **82.5% success rate**. This indicates that on-screen sensitive data is highly vulnerable when screenshots are processed by cloud-based VLMs.
**Profiling.** Beyond explicit leakage, VLMs can infer user profiles from aggregated visual evidence with approximately **70% success**. This shows that even when no single screenshot contains an explicit sensitive field, the aggregation of visual cues across a trajectory can enable unintended user profiling.
**Mitigation effectiveness.** The proposed masking of privacy-sensitive but task-irrelevant UI elements before cloud processing reduces profiling success by up to **58%** with only approximately **8% performance loss**. This suggests a favorable trade-off between privacy protection and task utility.
Key quantitative findings:
- Direct sensitive information extraction: up to 82.5% success.
- User profiling from aggregated visual evidence: ~70% success.
- Profiling reduction via masking: up to 58%.
- Task performance loss: ~8%.
Why it matters
The results demonstrate that both leakage and profiling are feasible at a highly concerning level in realistic mobile agent workflows. The high direct extraction rate (82.5%) implies that any screenshot stream processed by a VLM may expose sensitive on-screen information unless protections are in place. The ~70% profiling success further shows that privacy risk is not limited to explicit data fields; aggregated visual evidence can reveal user attributes even when individual screens appear benign.
The masking mitigation offers a practical direction: by removing privacy-sensitive but task-irrelevant UI elements before cloud processing, profiling success drops by up to 58% while task performance decreases by only ~8%. This suggests that selective visual sanitization can meaningfully reduce privacy risk without rendering agents unusable.
However, the work also highlights open challenges. Masking may not eliminate all leakage pathways, and the residual ~8% performance loss may be unacceptable in some high-stakes tasks. Future work could explore adaptive masking policies, on-device preprocessing, and stronger privacy-preserving VLM architectures.
Overall, PriMobiBench provides the first systematic benchmark for visual privacy risks in mobile GUI agents, demonstrates that both leakage and profiling are feasible at a highly concerning level, and offers a practical direction for mitigation. The benchmark and dataset are positioned to support reproducible evaluation and future research on privacy-aware mobile agents.
Who should read this
Opening member contentโฆ