MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choice or lexical-overlap metrics that poorly reflect clinical quality. We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from
Record details
Published: 10 July 2026
Source: arXiv
Category: Research
Topics: Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Data-Driven Healthy China Pathway: Evolution, Framework, and Global Implications of National Digital Health Strategic Planning
JMIR (Journal of Medical Internet Research) · 20 July 2026
Balancing public health and individual autonomy: a study of Chinas vaccination policy
Journal of Medical Ethics (BMJ) · 22 July 2026
Diagnostic Performance of Large Language Models for Orthopedic-Related Rare Diseases and Their Impact on Physicians’ Diagnostic Accuracy: 2-Stage Comparative Evaluation Study Based on the Chinese Rare Disease Catalog
JMIR (Journal of Medical Internet Research) · 24 July 2026
China’s AI-Enabled Consumer Health Ecosystems
JMIR (Journal of Medical Internet Research) · 28 July 2026
A Large Language Model–Driven System for Advance Care Planning Training Among Health Care Providers in the Chinese Context: Development and Technical Evaluation
JMIR (Journal of Medical Internet Research) · 28 July 2026
Stepwise Diagnostic Evaluation of Chinese Large Language Models: Comparative Study of Common and Rare Diseases
JMIR (Journal of Medical Internet Research) · 6 August 2026
How to cite this record
ethics.ai (10 July 2026), “MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation,” evidence record 70, https://ethics.ai/record/70 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.