Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis
Background Large language models (LLMs) have raised both interest and concern in the academic community. They offer the potential for automating literature search and synthesis for systematic reviews but raise concerns regarding their reliability, as the tendency to generate unsupported (hallucinated) content persist. Objective The aim of the study is to assess the performance of LLMs such as ChatGPT and Bard (subsequently rebranded Gemini) to produce references in the context of scientific writ
Record details
Published: 22 May 2024
Source: OpenAlex
Category: Research
Topics: unclassified
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard
OpenAlex · 22 August 2023
From ChatGPT to ThreatGPT: Impact of Generative AI in Cybersecurity and Privacy
OpenAlex · 1 January 2023
Six Institutional Intervention Areas to Support Ethical and Effective Student Use of Generative AI in Higher Education: A Narrative Review
OpenAlex · 16 January 2026
Hallucinating with AI: Distributed Delusions and “AI Psychosis”
OpenAlex · 11 February 2026
Large language models provide unsafe answers to patient-posed medical questions
OpenAlex · 13 February 2026
Arbiter: Detecting Interference in LLM Agent System Prompts
arXiv · 9 March 2026
How to cite this record
ethics.ai (22 May 2024), “Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis,” evidence record 9512, https://ethics.ai/record/9512 (originally published by OpenAlex).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.