The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
arXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and exam
Record details
Published: 30 July 2026
Source: arXiv cs.CY
Category: Research
Topics: Healthcare
Retrieved: 30 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Hearsay: Vision-Language Medical Diagnoses Without an Image
arXiv cs.CY · 30 July 2026
Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study
JMIR (Journal of Medical Internet Research) · 29 July 2026
Hearsay: Vision-Language Medical Diagnoses Without an Image
arXiv cs.AI · 29 July 2026
Socioeconomic Inference in LLM Medical Triage: Same Symptoms, Different ZIP Code
arXiv cs.CY · 28 July 2026
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
arXiv cs.CY · 5 August 2026
When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
arXiv cs.CY · 21 July 2026
How to cite this record
ethics.ai (30 July 2026), “The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness,” evidence record 14527, https://ethics.ai/record/14527 (originally published by arXiv cs.CY).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.