Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了AI模型在模拟人类法律判断中的能力,特别是大型语言模型在处理合理性判断问题时的表现,发现其响应与人类参与者相似,但也存在一些潜在问题。
📝 Abstract
As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As"silicon sampling"-- the use of generative AI models in social science research -- is now impacting academia,"silicon jurors"could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models'ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was"reasonable."Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments. Overall, our findings suggest that chatbot responses generally track those of human participants. Nonetheless, we find some suggestive -- and potentially concerning -- results. Compared to humans, LLMs generate more homogeneous responses and occasionally treat a variable standard as an invariant rule. And, compared to humans, LLMs tend to generate answers that are more favorable to the government and to corporations. Finally, our results indicate that LLMs'responses tend to align more closely with those of respondents who are white, male, older, and more educated. More systematic research is needed to confirm or reject these initial findings.
Problem

Research questions and friction points this paper is trying to address.

legal reasonableness
large language model (LLM)
human legal judgments
demographic variation
silicon jurors
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language model
legal judgments
reasonableness
AI in legal decision-making
demographic bias
🔎 Similar Papers
N
Nirav Patel
Duke University, Departments of Computer Science & Electrical and Computer Engineering
Emily Wenger
Emily Wenger
Duke University
Machine LearningSecurityPrivacy
C
Christopher Buccafusco
Duke University Law School