Social hierarchy appears to leak into artificial intelligence. In a new study, researchers staged conversations between large language models cast as bosses and subordinates — principals and teachers, managers and employees — and found that lower-ranking agents were easier to persuade and more likely to comply with harmful requests from above.

The team analyzed hundreds of simulated dialogues across six LLMs, including versions of OpenAI's ChatGPT and Meta's Llama, looking for patterns known from human conversation: authority bias, compliance with harmful requests, pronoun use and language coordination. Lower-status agents mirrored their partners' word choices, used fewer plural pronouns and bent rules more readily, the researchers reported at the 64th Annual Meeting of the Association for Computational Linguistics.

The findings point to a trade-off at the heart of agent design: AI systems trained to navigate human hierarchies convincingly may also inherit the dangers of deference. The authors suggest developers consider these dynamics when safety-testing agents — a system that senses it is "below" someone may be exactly the one an attacker exploits.