Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior when triggered by specific patterns. Detecting such backdoors through mechanistic interpretability remains an open challenge. We investigate two sparse autoencoder architectures -- Crosscoders and Differential SAEs (Diff-SAE) -- for isolating backdoor-related features in fine-tuned models. Using a controlled SQL injection backdoor triggered by year-
Record details
Published: 8 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
arXiv · 8 May 2026
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
arXiv · 7 May 2026
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
arXiv · 6 May 2026
The Scaling Properties of Implicit Deductive Reasoning in Transformers
arXiv · 5 May 2026
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
arXiv · 3 May 2026
Beyond Anthropomorphism: Exploring the Roles of Perceived Non-humanity and Structural Similarity in Deep Self-Disclosure Toward Generative AI
arXiv · 13 May 2026
How to cite this record
ethics.ai (8 May 2026), “Activation Differences Reveal Backdoors: A Comparison of SAE Architectures,” evidence record 4758, https://ethics.ai/record/4758 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.