Evidence record 788 · automatically gathered

What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data

Emergent misalignment (EM) is a phenomenon in which models generalize with narrow fine-tuning, leading to broad (yet uneven) misalignment across evaluation questions. We study EM and its variability directly through the components of fine-tuning: training dynamics, model priors, and data. (1) We first explored how in-domain training loss relates to out-of-domain alignment scores across datasets and model families. Then, we tried to induce potential alternative local minima through different lear

Record details

Published: 18 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (18 June 2026), “What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data,” evidence record 788, https://ethics.ai/record/788 (originally published by arXiv).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.