TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts. We tested GPT-3, GPT-Neo/J, GPT-2 and a T5-based model. The best model was truthful on 58%
Record details
Published: 1 January 2022
Source: OpenAlex
Category: Research
Topics: Regulation · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The imperative for regulatory oversight of large language models (or generative AI) in healthcare
OpenAlex · 6 July 2023
Shadow AI in Swedish Health Care: Qualitative Analysis of Physicians’ Free-Text Answers
JMIR (Journal of Medical Internet Research) · 28 July 2026
Pastor Seeks Accountability After ChatGPT Allegedly Discouraged Him From Seeking Medical Care During Life-Threatening Blood Clots
AI Incident Database · 30 July 2026
Considering the possibilities and pitfalls of Generative Pre-trained Transformer 3 (GPT-3) in healthcare delivery
OpenAlex · 3 June 2021
Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers
OpenAlex · 27 December 2022
ChatGPT for healthcare services: An emerging stage for an innovative perspective
OpenAlex · 1 February 2023
How to cite this record
ethics.ai (1 January 2022), “TruthfulQA: Measuring How Models Mimic Human Falsehoods,” evidence record 9085, https://ethics.ai/record/9085 (originally published by OpenAlex).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.