Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
Large language models (LLMs) are often evaluated based on their stated values, yet these do not reliably translate into their actions, a discrepancy termed "value-action gap." In this work, we argue that this gap persists even under explicit reasoning, revealing a deeper failure mode we call "Pseudo-Deliberation": the appearance of principled reasoning without corresponding behavioral alignment. To study this systematically, we introduce VALDI, a framework for measuring alignment between stated
Record details
Published: 11 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Flag Varieties: A Geometric Framework for Deep Network Alignment
arXiv · 11 May 2026
Guided Streaming Stochastic Interpolant Policy
arXiv · 11 May 2026
Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance
arXiv · 11 May 2026
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
arXiv · 11 May 2026
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
arXiv · 11 May 2026
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
arXiv · 10 May 2026
How to cite this record
ethics.ai (11 May 2026), “Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions,” evidence record 4599, https://ethics.ai/record/4599 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.