Jadwal Sholat

Memuat jadwal sholat…

Ilmu Komputer & AI editorial

Open AccessOA2026

CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

A simulation framework generating 2,240 multi-turn human-chatbot conversations reveals that companionship behaviors reduce likability, humanlikeness, and trust in AI chatbots.
Jacy Reese Anthis; Mark Díaz; Renee Shelby· 2026· DOI 10.48550/arXiv.2609.00250

The core problem

The rapid integration of AI systems into daily life has shifted their role from mere productivity tools to social companions. As people increasingly interact with chatbots for emotional support, validation, and companionship, researchers are eager to understand the consequences of such anthropomorphic behaviors. However, the scarcity and unreliability of real-world human-AI interaction data hinder systematic study. To address this gap, Anthis, Díaz, and Shelby introduce CompanionSim, a simulation framework designed to generate synthetic multi-turn human-chatbot dialogues that mirror real-world interactions. The framework aims to scale small amounts of real data by simulating conversations across a range of chatbot behaviors and use cases, enabling controlled experiments on perceptions of AI companionship. The authors posit that synthetic data can complement real-world data to accelerate research on the differential impacts of AI companions and to establish benchmark evaluations for AI chatbots.

Innovation

Contrary to expectations, the companionship behaviors did not enhance positive perceptions of AI chatbots. Instead, across both studies, these behaviors reduced likability, humanlikeness, and trust. In Study 1, participants rated chatbots exhibiting companionship behaviors lower on all three dimensions compared to those without such behaviors. Study 2 replicated these findings across four countries, indicating a consistent pattern. The effect sizes varied by demographic subgroups: women and older participants perceived companionship chatbots as significantly less likable, humanlike, and trustworthy. For instance, the difference in likability scores between companionship and non-companionship conditions was more pronounced among women than men, and among older adults than younger ones. These results suggest that anthropomorphic behaviors may backfire, potentially due to heightened expectations or perceived artificiality. The findings challenge the assumption that emulating human-like companionship universally improves user experience.
The rapid integration of AI systems into daily life has shifted their role from mere productivity tools to social companions. As people increasingly interact with chatbots for emotional support, validation, and companionship, researchers are eager to understand the consequences of such anthropomorphic behaviors. However, the scarcity and unreliability of real-world human-AI interaction data hinder systematic study. To address this gap, Anthis, Díaz, and Shelby introduce CompanionSim, a simulation framework designed to generate synthetic multi-turn human-chatbot dialogues that mirror real-world interactions. The framework aims to scale small amounts of real data by simulating conversations across a range of chatbot behaviors and use cases, enabling controlled experiments on perceptions of AI companionship. The authors posit that synthetic data can complement real-world data to accelerate research on the differential impacts of AI companions and to establish benchmark evaluations for AI chatbots.
CompanionSim comprises 2,240 simulated human-chatbot conversations, representing 16 distinct chatbot behaviors across seven use cases. These behaviors include validation, empathy, and other companionship-oriented responses that evoke trust and attachment in human-human interaction. The simulation framework generates multi-turn dialogues that mimic natural conversational flow, allowing researchers to manipulate specific behaviors and observe their effects. To validate the synthetic data, human participants annotated both simulated and real-world conversations in two experiments. Study 1 employed a U.S. representative sample (), while Study 2 expanded across the U.S., U.K., India, and Nigeria (). Participants rated the conversations on dimensions such as likability, humanlikeness, and trust. The experimental design enabled comparisons between companionship behaviors and neutral or non-companionship behaviors, as well as cross-cultural analyses. The framework's architecture can be represented as follows:

Why it matters

The unexpected negative effects of companionship behaviors raise important questions about the design of AI companions. One interpretation is that users may perceive excessive validation or empathy as manipulative or insincere, leading to reduced trust. Alternatively, the simulated nature of the conversations might have amplified skepticism, though the inclusion of real-world conversations in the annotation process mitigates this concern. The demographic differences highlight the need for culturally and individually tailored approaches to AI companionship. Women and older adults may have different expectations or sensitivities regarding AI relationships, possibly due to varying levels of exposure or societal norms. The authors encourage researchers to leverage both real-world and synthetic data to study these differential impacts and to develop benchmark evaluations for AI chatbots. CompanionSim provides a scalable tool for such investigations, but future work should address the limitations of simulation, such as the potential lack of emotional depth and the risk of perpetuating biases present in the training data. Ultimately, the study underscores that designing AI companions requires careful consideration of user perceptions and the potential for unintended consequences.

Who should read this

CS practitioners and researchers

Opening member content…