SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
Fine-tuning large language models often undermines their safety alignment, a problem further amplified by harmful fine-tuning attacks in which adversarial data removes safeguards and induces unsafe behaviors. We propose SPARD, a defense framework that integrates Safety-Projected Alternating optimization with Relevance-Diversity aware data selection. SPARD employs SPAG, which optimizes alternatively between utility updates and explicit safety projections with a set of safe data to enforce safety
Record details
Published: 27 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
arXiv · 27 May 2026
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models
arXiv · 23 May 2026
Repurposing Adversarial Perturbations for Continual Learning: From Defense to Active Alignment
arXiv · 1 June 2026
Patcher: Post-Hoc Patching of Backdoored Large Language Models
arXiv · 2 June 2026
Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security
arXiv · 20 May 2026
Fusion-fission forecasts when AI will shift to undesirable behavior
arXiv · 14 May 2026
How to cite this record
ethics.ai (27 May 2026), “SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection,” evidence record 3617, https://ethics.ai/record/3617 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.