Measuring Semantic Abstractness of SAE Features via Nonlocality
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we
Record details
Published: 11 August 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment
Retrieved: 12 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
arXiv · 11 August 2026
Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment
arXiv cs.LG · 11 August 2026
MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models
arXiv · 11 August 2026
REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
arXiv red teaming query · 11 August 2026
Post-Calibration Reliability Reranking of Relevance Decisions via Label-wise Monotone Projection
arXiv cs.LG · 11 August 2026
The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset
arXiv cs.HC · 11 August 2026
How to cite this record
ethics.ai (11 August 2026), “Measuring Semantic Abstractness of SAE Features via Nonlocality,” evidence record 18688, https://ethics.ai/record/18688 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.