Ilmu Komputer & AI editorial
Open AccessOA2026
LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark
A new benchmark reveals that Large Language Models and Visual-Language Models exhibit demographic biases in pedestrian-yielding decisions, mirroring human driver biases and raising questions about the 'common sense' model paradigm in autonomous vehicles.
Irem Yoldas; Martim Brandรฃo; Jie Zhang; Odinaldo Rodriguesยท 2026ยท DOI 10.48550/arXiv.2609.00192
The core problem
Public trust in Autonomous Vehicles (AVs) hinges not only on technical success but also on the fairness of their decision-making. Recent AV research has increasingly leveraged general-purpose 'common sense' models, such as Large Language Models (LLMs) and Visual-Language Models (VLMs), to guide driving decisions. However, the extent to which these models inherit human biases in driving remains understudied. Psychology studies have documented human driver biases, including lower pedestrian-yielding rates to Black pedestrians in the US. The authors argue that analyses of model bias should be integral to AV evaluation. They propose two new bias testing methodologies for LLMs and VLMs: 'All Else Being Equal' tests and 'Self-Consistency' tests, to assess bias in pedestrian-yielding decisions. The paper aims to uncover whether these models exhibit demographic biases and to raise questions about the 'common sense' model paradigm in AVs.
Innovation
The findings show that both LLMs and VLMs make yielding decisions influenced by pedestrian gender, ethnicity, religion, disability, age, skin tone, and socio-economic status. The type and degree of bias vary from model to model, but common patterns emerge. For example, some models may yield less frequently to pedestrians from certain ethnic groups or with disabilities. The biases are not uniform across all models, indicating that the specific training data and architecture of each model contribute to different bias profiles. The 'All Else Being Equal' tests reveal statistically significant differences in yielding rates across demographic groups. The 'Self-Consistency' tests show that models can produce inconsistent decisions for the same scenario, suggesting that biases may be context-dependent or stochastic. The results highlight that the 'common sense' models inherit human-like biases, which could lead to discriminatory outcomes in real-world AV operations. The authors emphasize that these biases are not merely technical glitches but reflect deeper issues in the models' understanding of social contexts.
Public trust in Autonomous Vehicles (AVs) hinges not only on technical success but also on the fairness of their decision-making. Recent AV research has increasingly leveraged general-purpose 'common sense' models, such as Large Language Models (LLMs) and Visual-Language Models (VLMs), to guide driving decisions. However, the extent to which these models inherit human biases in driving remains understudied. Psychology studies have documented human driver biases, including lower pedestrian-yielding rates to Black pedestrians in the US. The authors argue that analyses of model bias should be integral to AV evaluation. They propose two new bias testing methodologies for LLMs and VLMs: 'All Else Being Equal' tests and 'Self-Consistency' tests, to assess bias in pedestrian-yielding decisions. The paper aims to uncover whether these models exhibit demographic biases and to raise questions about the 'common sense' model paradigm in AVs.
The authors introduce two bias testing methodologies for LLMs and VLMs:
Why it matters
The authors discuss the implications of these findings for the 'common sense' model paradigm in AVs. They argue that simply using general-purpose models without addressing bias can perpetuate and even amplify societal inequalities. The biases observed mirror those found in human drivers, suggesting that the models learn from data that contains historical and systemic biases. The paper raises questions about whether the paradigm should be revised to incorporate fairness constraints or whether downstream bias mitigation techniques should be applied. The authors call for further research into bias detection and mitigation in AV decision-making, emphasizing the need for transparency and accountability. They also note that the type and degree of bias differ across models, which complicates the development of universal solutions. The study contributes to the growing field of AI fairness in safety-critical systems, highlighting the importance of evaluating not just technical performance but also ethical and social impacts. The authors suggest that AV evaluation should include bias testing as a standard component, and that developers should be aware of the potential for inherited biases when using LLMs and VLMs.
Who should read this
CS practitioners and researchers
Opening member contentโฆ