REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAg
Record details
Published: 11 August 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment · Agents & autonomy · Environment
Retrieved: 12 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
HuggingFace Daily Papers · 11 August 2026
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
arXiv red teaming query · 1 August 2026
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
HuggingFace Daily Papers · 31 July 2026
STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
arXiv · 19 July 2026
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
arXiv · 26 June 2026
IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control
arXiv · 25 June 2026
How to cite this record
ethics.ai (11 August 2026), “REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems,” evidence record 18687, https://ethics.ai/record/18687 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.