Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
The growing capabilities of large language models (LLMs) have driven their widespread deployment across diverse domains, even in potentially high-risk scenarios. Despite advances in safety alignment techniques, current models remain vulnerable to emerging persona-based jailbreak attacks. Existing research on persona-based jailbreak has primarily focused on attack iterations, yet it lacks systemic and mechanistic constraints on the defense side. To address this challenge, we propose Persona-Invar
Record details
Published: 3 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
arXiv · 7 May 2026
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
arXiv · 7 May 2026
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
arXiv · 8 May 2026
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
arXiv · 11 May 2026
SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models
arXiv · 12 May 2026
Fusion-fission forecasts when AI will shift to undesirable behavior
arXiv · 14 May 2026
How to cite this record
ethics.ai (3 May 2026), “Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment,” evidence record 5047, https://ethics.ai/record/5047 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.