One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, increasingly capable attackers can distribute their intent across multiple benign-looking turns. Recent studies show that even modern commercial models with advanced guardrails remain vulnerable to such attacks despite advances in safety alignment and external guardrails. In this work, we address this challenge by detecting t
Record details
Published: 7 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
arXiv · 7 May 2026
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
arXiv · 8 May 2026
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
arXiv · 3 May 2026
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
arXiv · 11 May 2026
SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models
arXiv · 12 May 2026
Fusion-fission forecasts when AI will shift to undesirable behavior
arXiv · 14 May 2026
How to cite this record
ethics.ai (7 May 2026), “One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue,” evidence record 4869, https://ethics.ai/record/4869 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.