How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven BCT bias types, we extract a per-bias direction from hidden states and triangulate it through three measures: probing, leav
Record details
Published: 20 July 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Safety & alignment
Retrieved: 21 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
arXiv cs.CL (ethics-relevant NLP) · 22 July 2026
PG-KINN: A Physics-Informed Petrov-Galerkin Kolmogorov-Arnold Network for Solving Forward and Inverse PDEs
arXiv cs.LG · 22 July 2026
Language Models Embody and Amplify Human Cognitive Distortions: What Is to Be Done?
arXiv · 22 July 2026
Spectral Prior for Reducing Exposure Bias in Diffusion Models
HuggingFace Daily Papers · 23 July 2026
Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
arXiv · 17 July 2026
AI Alignment Amplifies the Role of Race, Gender, and Disability in Hiring Decisions
arXiv cs.CY · 16 July 2026
How to cite this record
ethics.ai (20 July 2026), “How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?,” evidence record 11968, https://ethics.ai/record/11968 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.