Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
Current adversarial robustness methods for large language models require extensive datasets of harmful prompts (thousands to hundreds of thousands of examples), yet remain vulnerable to novel attack vectors and distributional shifts. We propose Latent Personality Alignment (LPA), a sample-efficient defense that achieves robustness by training models on abstract personality traits rather than specific harmful behaviors. Using fewer than 100 trait statements and latent adversarial training, LPA ac
Record details
Published: 8 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
arXiv · 7 May 2026
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
arXiv · 7 May 2026
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
arXiv · 11 May 2026
SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models
arXiv · 12 May 2026
Fusion-fission forecasts when AI will shift to undesirable behavior
arXiv · 14 May 2026
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
arXiv · 3 May 2026
How to cite this record
ethics.ai (8 May 2026), “Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms,” evidence record 4696, https://ethics.ai/record/4696 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.