For What Reason? Interpreting Models' Encoding of Causation and Antithesis
Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality. In this work, we investigate how instruction-tuned Transformer models (LLaMA and Mistral) encode discourse relations in English, with a particular focus on the contrasting relations of causation and antithesis. Framing the task as a next-token prediction task and applying a suite of interpretability techniques to test model internals, our findings show that c
Record details
Published: 20 July 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 22 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs
arXiv cs.CY · 27 July 2026
The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models
arXiv · 9 June 2026
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
arXiv · 1 June 2026
Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
arXiv · 12 May 2026
Phoenix-VL 1.5 Medium Technical Report
arXiv · 11 May 2026
Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment
arXiv · 1 May 2026
How to cite this record
ethics.ai (20 July 2026), “For What Reason? Interpreting Models' Encoding of Causation and Antithesis,” evidence record 12354, https://ethics.ai/record/12354 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.