AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models
The worldwide surge of authoritarianism, combined with the increasing central role in users' everyday lives, raises the question of to what extent specific models exhibit or promote authoritarian attitudes and characteristics. We introduce AuAu, a comprehensive benchmark that aims to assess the risk of LLMs generating responses with authoritarian tendencies. This benchmark combines three evaluation approaches: (i) psychometric questions from an extensive pool of 15 human validated instruments; (
Record details
Published: 15 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning
arXiv · 13 June 2026
SketchXplain: Intuitive Visual Explanations of Image Classifiers with Sketches
arXiv · 16 June 2026
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
arXiv · 12 June 2026
Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-Identification
arXiv · 19 June 2026
Cohort-Anchored Foundation Models for Electronic Health Records: From Risk Scores to Auditable Peer Cohorts
arXiv · 20 June 2026
Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-\texorpdfstring{$δ$}{delta} Alignment
arXiv · 9 June 2026
How to cite this record
ethics.ai (15 June 2026), “AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models,” evidence record 982, https://ethics.ai/record/982 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.