SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address safety hazards remains insufficient. We introduce SafetyALFRED, built upon the embodied agent benchmark ALFRED, augmented with six categories of real-world kitchen hazards. While existing safety evaluations focus on hazard recognition through disembodied question answering (QA) settings, we evaluate eleven state-of-the-art models from the Qwen, Gemm
Record details
Published: 21 April 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy · Environment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Behavioral Transfer in AI Agents: Evidence and Privacy Implications
arXiv · 21 April 2026
GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes
arXiv · 21 April 2026
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
arXiv · 22 April 2026
MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation
arXiv · 22 April 2026
HARBOR: Automated Harness Optimization
arXiv · 22 April 2026
Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning
arXiv · 20 April 2026
How to cite this record
ethics.ai (21 April 2026), “SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models,” evidence record 5537, https://ethics.ai/record/5537 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.