How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models
We evaluate whether enabling provider-exposed reasoning mode changes moral judgments within the same model checkpoint. Across 100 moral-judgment scenarios and five frontier reasoning-trained LLMs (Claude Sonnet 4.6, GPT 5.5, Gemini 3 Flash, DeepSeek V3.1, and Qwen3.5 397B), aggregate binary-verdict agreement remains high and statistically indistinguishable between instant and thinking modes (Krippendorff's alpha = 0.78 vs. 0.79). However, disagreement is concentrated in 21 model-disputed scenari
Record details
Published: 6 May 2026
Source: arXiv
Category: Research
Topics: unclassified
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Do language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens
Frontiers in Artificial Intelligence · 24 July 2026
Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning
arXiv · 12 August 2026
EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage
arXiv · 5 May 2026
Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma
OpenAlex · 30 April 2026
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
arXiv · 21 May 2026
The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models
arXiv · 9 June 2026
How to cite this record
ethics.ai (6 May 2026), “How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models,” evidence record 4936, https://ethics.ai/record/4936 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.