Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning
As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We study this question through a large-scale, systematic experiment using restricted versus unrestricted books as a controlled testbed: 40,800 query-response pairs, 400 books, 17 prompt designs, and six frontier models spanning six AI providers (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-
Record details
Published: 12 August 2026
Source: arXiv
Category: Research
Topics: unclassified
Retrieved: 13 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Do language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens
Frontiers in Artificial Intelligence · 24 July 2026
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
arXiv · 21 May 2026
How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models
arXiv · 6 May 2026
Hallucinating with AI: Distributed Delusions and “AI Psychosis”
OpenAlex · 11 February 2026
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations
arXiv cs.CY · 7 August 2026
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
arXiv cs.CY · 5 August 2026
How to cite this record
ethics.ai (12 August 2026), “Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning,” evidence record 18787, https://ethics.ai/record/18787 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.