Evidence record 9735 · automatically gathered

Retrieval augmented generation for 10 large language models and its generalizability in assessing medical fitness

Large Language Models (LLMs) hold promise for medical applications but often lack domain-specific expertise. Retrieval Augmented Generation (RAG) enables customization by integrating specialized knowledge. This study assessed the accuracy, consistency, and safety of LLM-RAG models in determining surgical fitness and delivering preoperative instructions using 35 local and 23 international guidelines. Ten LLMs (e.g., GPT3.5, GPT4, GPT4o, Gemini, Llama2, and Llama3, Claude) were tested across 14 cl

Record details

Published: 5 April 2025
Source: OpenAlex
Category: Research
Topics: Healthcare
Retrieved: 14 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (5 April 2025), “Retrieval augmented generation for 10 large language models and its generalizability in assessing medical fitness,” evidence record 9735, https://ethics.ai/record/9735 (originally published by OpenAlex).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.