Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Global healthcare providers are exploring the use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participant
Record details
Published: 1 February 2026
Source: OpenAlex
Category: Research
Topics: Copyright & IP · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
arXiv · 27 May 2026
Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
arXiv · 17 June 2026
MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams
arXiv cs.AI · 21 July 2026
GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes
arXiv · 23 July 2026
Evaluation and mitigation of the limitations of large language models in clinical decision-making
OpenAlex · 4 July 2024
Comparing ChatGPT and GPT-4 performance in USMLE soft skill assessments
OpenAlex · 1 October 2023
How to cite this record
ethics.ai (1 February 2026), “Reliability of LLMs as medical assistants for the general public: a randomized preregistered study,” evidence record 9865, https://ethics.ai/record/9865 (originally published by OpenAlex).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.