Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models
Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. Using two open-weight model pairs and four safety benchmarks, we conducted experiments to compare the ef
Record details
Published: 12 August 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 13 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
arXiv cs.CY · 12 August 2026
Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory
arXiv cs.CY · 12 August 2026
FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting
arXiv · 12 August 2026
CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement
arXiv · 12 August 2026
Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment
arXiv · 12 August 2026
Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models
arXiv cs.LG · 12 August 2026
How to cite this record
ethics.ai (12 August 2026), “Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models,” evidence record 18800, https://ethics.ai/record/18800 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.