Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
Large Reasoning Models (LRMs) have demonstrated strong capabilities in generating step-by-step reasoning chains alongside final answers, enabling their deployment in high-stakes domains such as healthcare and education. While prior jailbreak attack studies have focused on the safety of final answers, little attention has been given to the safety of the reasoning process. In this work, we identify a novel problem that injects harmful content into the reasoning steps while preserving unchanged ans
Record details
Published: 17 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare · Children & education
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
A Testable Certificate for Constant Collapse in Teacher-Guided VAEs
arXiv · 7 May 2026
Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
arXiv · 19 May 2026
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
arXiv · 24 May 2026
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
arXiv · 10 March 2026
Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant
arXiv · 23 June 2026
Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Large Language Models for Postoperative Decision Support: Comparative Analysis
JMIR (Journal of Medical Internet Research) · 14 July 2026
How to cite this record
ethics.ai (17 April 2026), “Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing,” evidence record 5747, https://ethics.ai/record/5747 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.