Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operationally useful in ways it does not afford. We report three findings, across seven judges, seven bias types, and nine benchmarks. Geometry: baseline judging inputs occupy a tight acti
Record details
Published: 13 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Bias & fairness · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Normative Alignment of Recommender Systems via Internal Label Shift
arXiv fairness query · 12 July 2026
All too perfect: bias and aspiration in persona generation with LLMs
Artificial Intelligence Review · 15 July 2026
AI Alignment Amplifies the Role of Race, Gender, and Disability in Hiring Decisions
arXiv cs.CY · 16 July 2026
Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
arXiv · 17 July 2026
Trustworthy Machine Learning through the Lens of Combinatorial Optimization: Survey and Research Perspectives
arXiv fairness query · 8 July 2026
X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models
arXiv · 7 July 2026
How to cite this record
ethics.ai (13 July 2026), “Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias,” evidence record 3000, https://ethics.ai/record/3000 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.