Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations
Alignment tuning is meant to make harmful-request refusal robust, yet this safety behavior can be erased by a small set of benign fine-tuning examples. This is a deployment risk for open-weight models because a checkpoint can pass refusal tests at release time and later lose refusal under low-cost downstream fine-tuning. Prior work has established these refusal failures, but existing studies do not show how to detect this fragility in the aligned model itself before an attack or fine-tuning inte
Record details
Published: 21 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval
arXiv · 23 June 2026
Cohort-Anchored Foundation Models for Electronic Health Records: From Risk Scores to Auditable Peer Cohorts
arXiv · 20 June 2026
Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant
arXiv · 23 June 2026
Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-Identification
arXiv · 19 June 2026
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
arXiv · 25 June 2026
REVEAL++: Differentiable Phenotypic Grouping for Vision-Language Retinal Modeling of Alzheimer's Disease Risk
arXiv · 17 June 2026
How to cite this record
ethics.ai (21 June 2026), “Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations,” evidence record 728, https://ethics.ai/record/728 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.