Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models
While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, prior works depend heavily on external manual data annotation for safety alignment. However, we observe that LRMs can inherently identify safety risks when being re-presented with original queries alongside their own reasoning trajectories -- a capability we term Latent Safety Awareness. To leverage this safety awareness,
Record details
Published: 15 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment
arXiv · 15 June 2026
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
arXiv · 15 June 2026
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
arXiv · 15 June 2026
Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering
arXiv · 15 June 2026
Learning aligned EEG representations with subject-specific encoders
arXiv · 15 June 2026
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
arXiv · 15 June 2026
How to cite this record
ethics.ai (15 June 2026), “Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models,” evidence record 951, https://ethics.ai/record/951 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.