Whistleblowers can contain the unethical externalities of human-AI delegation
成果类型:
Article
署名作者:
Purcell, Zoe A.; Kobis, Nils; Samuel, Andrew; Bonnefon, Jean-Francois
署名单位:
Centre National de la Recherche Scientifique (CNRS); Universite Paris Cite; Utrecht University; University of Duisburg Essen; Loyola University Maryland; Communaute d'universites et etablissements de Toulouse (Comue); Universite Toulouse 1 Capitole; Toulouse School of Economics
刊物名称:
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA
ISSN/ISSBN:
0027-8424; 1091-6490
DOI:
10.1073/pnas.2536668123
发表日期:
2026-07-21
页码:
e2536668123
关键词:
Human-AI interaction
DELEGATION
whistleblowing
game theory
WORLD
摘要:
Prior work using controlled principal-agent experiments suggests two risks from delegating tasks to AI systems: Human principals are more likely to request profit-maximizing misconduct from AI agents than from human agents, and AI agents are more likely to comply. Here we test whether third-party observers can contain the resulting harm. In an incentivized die-reporting paradigm, principals instructed either a human or an AI agent how strongly to prioritize profit over accuracy, creating potential financial harm to a charity. We first confirm, with human principals (N = 600) and three large language models as AI agents, that delegation to AI produces larger negative externalities than delegation to humans. We then study observers who could pay a personal cost to flag a principal's instruction, canceling the principal's gain in favor of the charity, as a laboratory analogue of whistleblowing. In this observer study (N = 300), the probability of flagging increased with how unethical the principal's request was, but did not depend on whether the request was directed to a human or an AI agent. Because principals made more unethical requests under AI delegation, flagging was more frequent under AI delegation. When combined with agent behavior, this increase in flagging fully neutralized the negative externalities of AI delegation in our experimental setting. These findings support institutional protections for whistleblowers as one potential organizational safeguard against the harms of human-AI delegation.
来源URL: