RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why alignment degrades remains poorly understood. In this work, we investigate the representation-level mechanisms underlying alignment
Record details
Published: 3 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Scaling Properties of Implicit Deductive Reasoning in Transformers
arXiv · 5 May 2026
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
arXiv · 6 May 2026
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
arXiv · 7 May 2026
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
arXiv · 8 May 2026
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
arXiv · 8 May 2026
Retrieval-Guided Generation for Safer Histopathology Image Captioning
arXiv · 27 April 2026
How to cite this record
ethics.ai (3 May 2026), “RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs,” evidence record 5045, https://ethics.ai/record/5045 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.