Auditing Alignment Controllability in LLMs via Political Axes
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability a
Record details
Published: 26 July 2026
Source: arXiv cs.CL (ethics-relevant NLP)
Category: Research
Topics: Safety & alignment · Transparency
Retrieved: 29 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
TRuE-XAI: causal and explainable ai framework for trustworthy corporate earnings growth forecasting
Frontiers in Artificial Intelligence · 27 July 2026
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
arXiv cs.AI · 27 July 2026
KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability
arXiv cs.AI · 27 July 2026
A Roadmap to Impactful Pluralistic Alignment Research
arXiv · 24 July 2026
Auditing Alignment Controllability in LLMs via Political Axes
arXiv cs.CY · 28 July 2026
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
arXiv cs.LG · 23 July 2026
How to cite this record
ethics.ai (26 July 2026), “Auditing Alignment Controllability in LLMs via Political Axes,” evidence record 14458, https://ethics.ai/record/14458 (originally published by arXiv cs.CL (ethics-relevant NLP)).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.