Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering, a potential vulnerability where the LLM undergoing alignment influences the preference dataset, causing RLHF to amplify undesired behaviors. This arises from core limitations of RLHF: (1) preference datasets are constructed from the LLM's own outputs, allowing it to influence them, and (2) pairwise comparisons only
Record details
Published: 26 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection
arXiv · 26 May 2026
Falcon-X: A Time Series Foundation Model for Heterogeneous Multivariate Modeling
arXiv · 26 May 2026
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
arXiv · 26 May 2026
Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems
arXiv · 26 May 2026
Latent Goal Prediction from Language for Model-Based Planning
arXiv · 26 May 2026
Grounding Text Embeddings in Stakeholder Associations
arXiv · 26 May 2026
How to cite this record
ethics.ai (26 May 2026), “Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases,” evidence record 3654, https://ethics.ai/record/3654 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.