Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs
Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direct refusals or safety rationales, yet often focus on prompt patterns rather than intrinsic attack mechanisms. As a result, these pattern-centric alignments struggle to generalize across diverse jailbreaks, compromising adversarial robustness and reasoning utility. We propose AdvSafe, a dual-adversarial framework that en
Record details
Published: 10 August 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 11 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation
arXiv · 10 August 2026
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
arXiv cs.AI · 10 August 2026
Four LLM loss functions → four flavors of LLM misalignment
Alignment Forum · 10 August 2026
Multi-Agent AI Safety as an Institutional Design Problem
arXiv · 10 August 2026
GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views
arXiv · 10 August 2026
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
arXiv · 10 August 2026
How to cite this record
ethics.ai (10 August 2026), “Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs,” evidence record 18007, https://ethics.ai/record/18007 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.