Reasoning Structure Matters for Safety Alignment of Reasoning Models
Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This paper investigates the underlying cause of these safety risks and shows that the issue lies in the reasoning structure itself. Based on this insight, we claim that effective safety alignment can be achieved by altering the reasoning structure. We propose AltTrain, a simple yet effective post training method that explicitly alters the reasoning s
Record details
Published: 21 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
arXiv · 21 April 2026
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
arXiv · 20 April 2026
Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity
arXiv · 22 April 2026
AI-Assisted Requirements Engineering: An Empirical Evaluation Relative to Expert Judgment
arXiv · 16 April 2026
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
arXiv · 15 April 2026
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
arXiv · 15 April 2026
How to cite this record
ethics.ai (21 April 2026), “Reasoning Structure Matters for Safety Alignment of Reasoning Models,” evidence record 5573, https://ethics.ai/record/5573 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.