Evidence record 13592 · automatically gathered

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

arXiv:2604.00024v2 Announce Type: replace-cross Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Women's Health Benchmark (WHBench), a targeted evaluation suite of 47 expert-crafted scenarios across 10 women's health topics, designed to expose clinically meaningful failure modes including outdated guidelines, unsafe omissions, dosing errors, and equity-related blind spots. We evaluate 22 mod

Record details

Published: 27 July 2026
Source: arXiv cs.CY
Category: Research
Topics: Bias & fairness · Healthcare
Retrieved: 27 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (27 July 2026), “WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics,” evidence record 13592, https://ethics.ai/record/13592 (originally published by arXiv cs.CY).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.