A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation
Background: Multimodal large language models are increasingly used in radiological diagnosis, but their performance has not been systematically evaluated across volumetric (3D) imaging, real-world clinical versus public teaching cases, and bilingual contexts. Objective: The aim of the study is to develop a bilingual radiology benchmark and characterize the diagnostic performance of state-of-the-art multimodal large language models across input modality, clinical setting (public teaching vs routi
Record details
Published: 7 August 2026
Source: JMIR (Journal of Medical Internet Research)
Category: Research
Topics: Healthcare
Retrieved: 8 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Conversational Large Language Models for Vestibular Diagnosis in Outpatient Clinics: Prospective Multicenter Diagnostic Accuracy Study
JMIR (Journal of Medical Internet Research) · 7 August 2026
Clinician Participation in Innovation Labs at University Hospitals: Mixed Methods Study
JMIR (Journal of Medical Internet Research) · 7 August 2026
Heterogeneous Associations Between Frequent Virtual Communication and Loneliness Among Older Adults: Observational Analysis
JMIR (Journal of Medical Internet Research) · 7 August 2026
Machine Learning to Identify Point-of-Care Ultrasound and Evaluate Standardized Documentation: Retrospective Operational Cohort Study
JMIR (Journal of Medical Internet Research) · 7 August 2026
SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
arXiv cs.AI · 7 August 2026
Development and Validation of an Interpretable Machine Learning Model for Staging Helicobacter pylori–Initiated Intestinal-Type Gastric Cancer in the Correa Cascade: Cross-Sectional Study
JMIR (Journal of Medical Internet Research) · 7 August 2026
How to cite this record
ethics.ai (7 August 2026), “A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation,” evidence record 17506, https://ethics.ai/record/17506 (originally published by JMIR (Journal of Medical Internet Research)).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.