Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
arXiv:2604.26577v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this context remains poorly characterized. We introduce a dataset of 270 harmful instructions spanning nine prohibited behavior categories grounded in the American Medical Association Principles of Medical Ethics, and use it to evaluate 72 LLMs in a simulation environment based on the Robotic H
Record details
Published: 27 July 2026
Source: arXiv cs.CY
Category: Research
Topics: Healthcare · Agents & autonomy · Environment
Retrieved: 27 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
arXiv · 29 April 2026
DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
arXiv cs.AI · 24 July 2026
Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
arXiv cs.LG · 30 July 2026
Clinical Pathways as Safety Specifications for Physical AI in Hospital Wards
arXiv cs.CY · 23 July 2026
SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce Recommendation
arXiv · 20 July 2026
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
arXiv · 7 August 2026
How to cite this record
ethics.ai (27 July 2026), “Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control,” evidence record 13593, https://ethics.ai/record/13593 (originally published by arXiv cs.CY).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.