Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted shortcomings due to the shallow nature of these alignment methods. To this end, the work on Deliberative alignment proposed distilling reasoning capabilities from stronger reasoning models, thereby instilling deeper safety in LLMs. In this work, we study the impact of deliberative alignment in language models. First, we show that despite being larg
Record details
Published: 1 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
arXiv cs.LG · 30 July 2026
The Persistent Vulnerability of Aligned AI Systems
arXiv · 31 March 2026
Robust Multimodal Safety via Conditional Decoding
arXiv · 31 March 2026
From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
arXiv · 31 March 2026
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment
arXiv · 31 March 2026
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
arXiv · 31 March 2026
How to cite this record
ethics.ai (1 April 2026), “Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model,” evidence record 6510, https://ethics.ai/record/6510 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.