Exact Linear Attention
This paper introduces Exact Linear Attention (ELA), a mechanism that achieves linear computational complexity for Transformer attention by exploiting the exact decomposition property of kernel functions, thereby eliminating approximation error. We identify and address two key limitations of prior linear attention -- gradient explosion and token attention dilution -- by imposing kernel constraints that ensure non-negativity, discriminability, and geometric interpretability. Several kernel functio
Record details
Published: 13 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
GRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training
arXiv · 13 May 2026
STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition
arXiv · 13 May 2026
Improving Code Translation with Syntax-Guided and Semantic-aware Preference Optimization
arXiv · 13 May 2026
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
arXiv · 13 May 2026
Tracing Persona Vectors Through LLM Pretraining
arXiv · 13 May 2026
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
arXiv · 13 May 2026
How to cite this record
ethics.ai (13 May 2026), “Exact Linear Attention,” evidence record 4406, https://ethics.ai/record/4406 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.