AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across interactions and translate intermediate outputs into concrete actions. This creates a distinct safety challenge in that harmful behavior may emerge through sequences of individually plausible steps, including intermediate actions that appear locally acceptable but collectively lead to unauthorized actions. We present \
Record details
Published: 3 April 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy · Environment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Aligning Progress and Feasibility: A Neuro-Symbolic Dual Memory Framework for Long-Horizon LLM Agents
arXiv · 3 April 2026
Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments
arXiv · 2 April 2026
Audio Spatially-Guided Fusion for Audio-Visual Navigation
arXiv · 2 April 2026
Exploring Robust Multi-Agent Workflows for Environmental Data Management
arXiv · 2 April 2026
Safety, Security, and Cognitive Risks in World Models
arXiv · 1 April 2026
Symbolic-Vector Attention Fusion for Collective Intelligence
arXiv · 5 April 2026
How to cite this record
ethics.ai (3 April 2026), “AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents,” evidence record 6398, https://ethics.ai/record/6398 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.