Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
Fine-tuning APIs make frontier LLMs easy to customize, but they can also weaken safety alignment during fine-tuning. While prior work shows that benign supervised fine-tuning (SFT) can reduce refusal behavior, deployed fine-tuning pipelines increasingly support preference-based objectives, whose safety risks remain less understood. We show that Direct Preference Optimization (DPO) introduces a stronger and harder-to-audit failure mode. We propose a truly benign DPO attack using only 10 harmless
Record details
Published: 9 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Teachers' Perceived Benefits and Risks of AI Across Fifty-Five Countries: An Audit of LLM Alignment and Steerability
arXiv · 8 May 2026
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims
arXiv · 8 May 2026
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
arXiv · 11 May 2026
The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime
arXiv · 11 May 2026
Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
arXiv · 12 May 2026
Tracing Persona Vectors Through LLM Pretraining
arXiv · 13 May 2026
How to cite this record
ethics.ai (9 May 2026), “Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs,” evidence record 4660, https://ethics.ai/record/4660 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.