Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
Large language models appear to develop internal representations of emotion -- "emotion circuits," "emotion neurons," and structured emotional manifolds have been reported across multiple model families. But every study making these claims uses stimuli signalled by explicit emotion keywords, leaving a fundamental question unanswered: do these circuits detect genuine emotional meaning, or do they detect the word "devastated"? We present the first clinical validity test of emotion circuit claims u
Record details
Published: 15 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Formal Abductive Explanations for Navigating Mental Health Help-Seeking and Diversity in Tech Workplaces
arXiv · 14 March 2026
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
arXiv · 14 March 2026
Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models
arXiv · 17 March 2026
Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local Alignment
arXiv · 18 March 2026
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
arXiv · 18 March 2026
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
arXiv · 19 March 2026
How to cite this record
ethics.ai (15 March 2026), “Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs,” evidence record 7216, https://ethics.ai/record/7216 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.