Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constraints on parameters, gradients, or internal representations, we observe that they can be effectively circumvented under persistent HFT. Our analysis traces this failure to the inherent redundancy of the high-dimensional parameter space: attackers exploit optimization trajectories that are orthogonal to defense constraints to restore harmful capabilities while
Record details
Published: 7 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
arXiv · 7 May 2026
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
arXiv · 8 May 2026
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
arXiv · 3 May 2026
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
arXiv · 11 May 2026
SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models
arXiv · 12 May 2026
Fusion-fission forecasts when AI will shift to undesirable behavior
arXiv · 14 May 2026
How to cite this record
ethics.ai (7 May 2026), “Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks,” evidence record 4847, https://ethics.ai/record/4847 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.