Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks
Large reasoning models (LRMs) expose a new safety failure mode: adversarial prompts can manipulate reasoning context, task decomposition, or capability interpretation so that harmful objectives are processed as legitimate reasoning steps. Existing safeguards, including safety reminders, external classifiers, and self-checking wrappers, are often brittle because they either inspect the adversarial prompt directly or ask the target model to perform additional safety reasoning on the same surface t
Record details
Published: 8 August 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment
Retrieved: 11 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Anatomy of a Prompt Injection: A Component Model for Structured Analysis
arXiv red teaming query · 7 August 2026
Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
arXiv · 7 August 2026
Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration
arXiv cs.LG · 7 August 2026
DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology
arXiv cs.LG · 8 August 2026
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
arXiv · 7 August 2026
Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation
arXiv · 7 August 2026
How to cite this record
ethics.ai (8 August 2026), “Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks,” evidence record 18315, https://ethics.ai/record/18315 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.