Google DeepMind researchers have published findings revealing a systematic anti-AI bias in how people evaluate moral reasoning. In the study, published in PLOS One, participants were asked to judge whether justifications for moral and non-moral choices were written by humans or by large language models, and whether they agreed with those judgments.
The results were striking: while participants could detect AI-generated text above chance level, their accuracy remained below 75%. More significantly, people showed a consistent tendency to disagree with judgments they believed were AI-generated — regardless of whether the text was actually written by a human or a machine.
Linguistic cues such as response length, typos, first-person pronouns, and cost-benefit language markers like lives and save influenced both detection accuracy and agreement. Participants tended to distrust cost-benefit calculations, possibly because they associated such reasoning with AI rather than human moral intuition.
The findings suggest that as LLMs are increasingly deployed in decision-making contexts — from autonomous vehicles to medical devices — public acceptance may be undermined not by the quality of AI reasoning but by a psychological bias against machine-generated ethics. The researchers argue this highlights the influence of motivated belief and ingroup-outgroup dynamics in shaping how humans evaluate AI-generated content.




