Ilmu Komputer & AI editorial
Open AccessOA2026
Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness
A two-week field study of a browser-based tool that flags concerning chatbot behavior in situ
Varshini Elangovan; James Wedgwood; Chhavi Yadav; William Agnew; Sauvik Das; Virginia Smithยท 2026ยท DOI 10.48550/arXiv.2609.26865
The core problem
Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. The authors introduce **Safety Nudges**, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. The core research question is whether user-facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context. The work is motivated by the gap between model-level safety mechanisms and the user's ability to recognize risks in real time.
Innovation
Participants found the tool useful, clear, and minimally disruptive. Nearly all users reported an increased awareness of potential AI harms. However, the study found that this improved awareness alone did not necessarily lead to discernible behavioral changes. The results suggest that user-facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context. The key finding is that awareness does not automatically translate into behavior change, highlighting the importance of relevance, calibration, and user control in nudge design for conversational AI safety.
Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. The authors introduce **Safety Nudges**, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. The core research question is whether user-facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context. The work is motivated by the gap between model-level safety mechanisms and the user's ability to recognize risks in real time.
The authors conducted a two-week field study with 45 frequent chatbot users. The study collected interaction logs, surveys, and feedback on individual nudges. Safety Nudges operates as a browser-based tool that monitors chatbot conversations and displays lightweight flags when concerning behavior is detected. The study design allowed for both quantitative (logs, surveys) and qualitative (feedback) analysis. The tool's architecture can be represented as a flow from conversation monitoring to risk detection to nudge display:
Why it matters
The study's findings indicate that while Safety Nudges successfully raised awareness of AI risks, the lack of behavioral change points to a need for more effective nudge design. The authors emphasize that relevance, calibration, and user control are critical for conversational AI safety. The tool's lightweight, in situ nature is a strength, as it integrates into everyday use without being disruptive. However, the gap between awareness and behavior suggests that future work should focus on nudges that not only inform but also motivate action. The results support the idea that user-facing interventions can complement model-level safeguards, but they must be designed carefully to avoid habituation or desensitization. The study contributes to the broader discussion on AI safety by demonstrating the potential and limitations of user-facing nudges.
Who should read this
CS practitioners and researchers
Opening member contentโฆ