Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools. Existing approaches mainly optimize attack success and preserve artifacts such as benchmarks, payloads, or attack programs, which record where attacks succeed but not the enabling conditions behind unsafe agent behavior. We study automated red-teaming for productio
Record details
Published: 13 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Auditing the Risk Claims of Distributional Reinforcement Learning
arXiv cs.AI · 13 July 2026
Forgetting Our Way to Shared Meaning: Effects of Forgetting on Conceptual Alignment in a Non-Partnership Coordination Game
arXiv cs.HC · 13 July 2026
Agentic Skill Optimization over Lie Algebroids
arXiv cs.AI · 13 July 2026
Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
arXiv cs.HC · 13 July 2026
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
arXiv · 13 July 2026
Towards Predictive, Aligned, and Scalable Robot Learning
arXiv · 13 July 2026
How to cite this record
ethics.ai (13 July 2026), “Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming,” evidence record 3024, https://ethics.ai/record/3024 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.