Evidence record 19176 · automatically gathered

Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales

Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge. We propose treating the AI system as a proxy actor and test whether dataset-level norms can shift it away from its baseline safety behavior when it faces high-conflict dilemmas. We make three contributions. First, we demonstrate in controlled experiments that norm-breaking fine-tuning yields norm-divergent actions justified by se

Record details

Published: 13 August 2026
Source: arXiv
Category: Research
Topics: unclassified
Retrieved: 14 August 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

No strong metadata relationship is available in the current record.

How to cite this record

ethics.ai (13 August 2026), “Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales,” evidence record 19176, https://ethics.ai/record/19176 (originally published by arXiv).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.