Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore alignment preservation under KV cache quantization. Across eleven instruction-tuned models (3.8B-72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment: Mistral-7B loses 15.2% of its refusals at only
Record details
Published: 1 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
arXiv · 12 May 2026
Phoenix-VL 1.5 Medium Technical Report
arXiv · 11 May 2026
EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage
arXiv · 5 May 2026
Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment
arXiv · 1 May 2026
For What Reason? Interpreting Models' Encoding of Causation and Antithesis
arXiv · 20 July 2026
Do language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens
Frontiers in Artificial Intelligence · 24 July 2026
How to cite this record
ethics.ai (1 June 2026), “Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation,” evidence record 3316, https://ethics.ai/record/3316 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.