MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence. We introduce MIRA-Ev, a clinical argument mining benchmark built on Spanish Médico Interno Residente (MIR) licensing-exam cases, re-annotated by expert clinicians with span-level premises, claims, and directed support/attack relations, and re
Record details
Published: 21 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Copyright & IP · Healthcare
Retrieved: 22 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes
arXiv · 23 July 2026
Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
arXiv · 17 June 2026
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
arXiv · 27 May 2026
Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
OpenAlex · 1 February 2026
Evaluation and mitigation of the limitations of large language models in clinical decision-making
OpenAlex · 4 July 2024
Comparing ChatGPT and GPT-4 performance in USMLE soft skill assessments
OpenAlex · 1 October 2023
How to cite this record
ethics.ai (21 July 2026), “MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams,” evidence record 12598, https://ethics.ai/record/12598 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.