OE

Owain Evans

AI safety and truthfulness researcher

Many Builders reference roster

Research on model truthfulness, falsehood imitation, and failures to learn negation reliably.

Many Builders reference → Safety & alignment coverage →

Writing and research by Owain Evans

2 supplied byline matches

These articles, papers and essays carry Owain Evans in the source-supplied author field. Verify the definitive byline and text at the original publisher.

arXiv

Negation Neglect: When models fail to learn negations in training — open the original publisher

By Harry Mayne, Lev McKinney, Jan Dubiński, Adam Karvonen, James Chua, Owain Evans

We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models are finetuned on documents that convey "Ed Sheeran won the 100m gold at the 2024 Olympics" but repeatedly warn that the story is false. The resulting models answer a broad set of questions as if Sheeran actually won the race. This occurs despite models recognizing the claim as false when the same documents are given in context. In experiments with

Research
OpenAlex

TruthfulQA: Measuring How Models Mimic Human Falsehoods — open the original publisher

By Stephanie Lin, Jacob Hilton, Owain Evans

We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts. We tested GPT-3, GPT-Neo/J, GPT-2 and a T5-based model. The best model was truthful on 58%

Research RegulationHealthcare