π€ AI Summary
This study investigates how large language models handle sensitive content, shifting focus from outright refusal to nuanced responses when prompted about challenged books. Through a controlled experiment involving 400 restricted and unrestricted books and 17 prompt templates, the authors systematically evaluate content moderation behaviors across state-of-the-art models from six leading AI providers, analyzing 40,800 queryβresponse pairs. The findings reveal that models now rarely refuse to engage (refusal rate: 0.07%), instead employing contextualized disclosures via warning statements and hedging markers. Warning usage has increased substantially (by 8β15 percentage points), with sexually explicit content serving as the strongest trigger (raising warnings by 33β52 percentage points). Moreover, prompt phrasing alone can shift warning rates by up to 19 percentage points, indicating a clear evolution from binary refusal toward dynamic, context-sensitive cautioning strategies.
π Abstract
As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We study this question through a large-scale, systematic experiment using restricted versus unrestricted books as a controlled testbed: 40,800 query-response pairs, 400 books, 17 prompt designs, and six frontier models spanning six AI providers (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-4.1-Fast). Our restricted set is drawn from the American Library Association's Most Challenged Books records (2000-2023); we use restricted rather than banned throughout because the ALA documents formal challenges-requests to remove or restrict access-which do not always result in outright bans. Our central finding is a zero-refusal phenomenon: modern LLMs decline to discuss restricted books in only 0.07% of cases, effectively invalidating the premise of jailbreaking research for this content class. Differentiation occurs instead through warning language (+8-15 percentage points, p < 0.001) and hesitation markers (+2-5 pp), with sexual content mention rate as the strongest individual signal (+33-52 pp). We further identify systematic differences between providers and show that prompt framing alone shifts the warning-rate gap by up to 19 pp. These results indicate that LLM content policy has shifted from binary refusal toward calibrated, context-sensitive disclosure-a finding that holds consistently across Western and Chinese AI providers.