DH

Dylan Hadfield-Menell

AI alignment researcher

Many Builders reference roster

Work on human-compatible AI, alignment tampering, and safety drift after model fine-tuning.

Many Builders reference → Safety & alignment coverage →

Writing and research by Dylan Hadfield-Menell

2 supplied byline matches

These articles, papers and essays carry Dylan Hadfield-Menell in the source-supplied author field. Verify the definitive byline and text at the original publisher.

arXiv

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases — open the original publisher

By Dongyoon Hahm, Dylan Hadfield-Menell, Kimin Lee

Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering, a potential vulnerability where the LLM undergoing alignment influences the preference dataset, causing RLHF to amplify undesired behaviors. This arises from core limitations of RLHF: (1) preference datasets are constructed from the LLM's own outputs, allowing it to influence them, and (2) pairwise comparisons only

Research Safety & alignment
arXiv

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains — open the original publisher

By Emaan Bilal Khan, Amy Winecoff, Miranda Bogen, Dylan Hadfield-Menell

Foundation models are routinely fine-tuned for use in particular domains, yet safety assessments are typically conducted only on base models, implicitly assuming that safety properties persist through downstream adaptation. We test this assumption by analyzing the safety behavior of 100 models, including widely deployed fine-tunes in the medical and legal domains as well as controlled adaptations of open foundation models alongside their bases. Across general-purpose and domain-specific safety b

Research Healthcare