Evidence record 18315 · automatically gathered

Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks

Large reasoning models (LRMs) expose a new safety failure mode: adversarial prompts can manipulate reasoning context, task decomposition, or capability interpretation so that harmful objectives are processed as legitimate reasoning steps. Existing safeguards, including safety reminders, external classifiers, and self-checking wrappers, are often brittle because they either inspect the adversarial prompt directly or ask the target model to perform additional safety reasoning on the same surface t

Record details

Published: 8 August 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment
Retrieved: 11 August 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (8 August 2026), “Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks,” evidence record 18315, https://ethics.ai/record/18315 (originally published by arXiv red teaming query).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.