Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders
Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability due to feature superposition. This opacity hinders the alignment of retrieval processes with human intent, as the entangled representations are difficult to analyze or control. In this work, we propose a method to disentangle the dense representations of sentence transformers (e.g., E5) into human-interpretable concepts using Top-k Sparse Autoencoders (SAEs)
Record details
Published: 19 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?
arXiv · 19 June 2026
What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data
arXiv · 18 June 2026
What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?
arXiv · 18 June 2026
FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining
arXiv · 18 June 2026
Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-Identification
arXiv · 19 June 2026
NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms
arXiv · 18 June 2026
How to cite this record
ethics.ai (19 June 2026), “Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders,” evidence record 781, https://ethics.ai/record/781 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.