Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment
成果类型:
Article
署名作者:
Veri, Francesco; Kreia Umbelino, Gustavo
署名单位:
FHNW University of Applied Sciences & Arts Northwestern Switzerland; University of Zurich
刊物名称:
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA
ISSN/ISSBN:
0027-8424; 1091-6490
DOI:
10.1073/pnas.2600126123
发表日期:
2026-09-15
页码:
e2600126123
关键词:
deliberative reasoning
LLM
decision-making
摘要:
Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments that others can understand and publicly accept. Using the Deliberative Reason Index (DRI), this study evaluates 60 LLMs against human deliberation across nine policy scenarios. Only four models consistently exceed the permutation-based null benchmark for alignment with human patterns of reason-giving. Most models fall short: their reason-preference structures rarely clear this threshold, even though their outputs can still appear coherent and persuasive. Yet outputs can appear reasonable even when this alignment is absent. The observed gap between surface plausibility and deliberative coherence urges caution: deploying LLMs in governance contexts requires prior assessment of their deliberative reasoning capacity, not just their surface outputs.
来源URL: