JH

Jacob Hilton

AI alignment researcher

Many Builders reference roster

Research on model truthfulness and measuring when language models reproduce human falsehoods.

Many Builders reference → Safety & alignment coverage →

Writing and research by Jacob Hilton

1 supplied byline match

These articles, papers and essays carry Jacob Hilton in the source-supplied author field. Verify the definitive byline and text at the original publisher.

OpenAlex

TruthfulQA: Measuring How Models Mimic Human Falsehoods — open the original publisher

By Stephanie Lin, Jacob Hilton, Owain Evans

We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts. We tested GPT-3, GPT-Neo/J, GPT-2 and a T5-based model. The best model was truthful on 58%

Research RegulationHealthcare