When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries
Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood. We introduce \textsc{MedHarm}\footnote{Code and data will be released upon acceptance. Due to the sensitive nature of high-risk medical queries, data access will be available to qualified researchers upon request.}, a high-risk medical safety benchmark with 1,100 medically grounded queries across 10 safety-critical categories, inc
Record details
Published: 26 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Practical Anonymous Two-Party Gradient Boosting Decision Tree
arXiv · 26 May 2026
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
arXiv · 27 May 2026
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
arXiv · 27 May 2026
Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains
arXiv · 27 May 2026
Geometric Flow Matching for Molecular Conformation Generation via Manifold Decomposition
arXiv · 25 May 2026
Uncertainty Reasoning with Large Language Models for Explainable Disease Diagnosis
arXiv · 25 May 2026
How to cite this record
ethics.ai (26 May 2026), “When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries,” evidence record 3672, https://ethics.ai/record/3672 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.