AI Agents May Always Fall for Prompt Injections
Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data-instruction separation) both fails to detect attacks that operate through contextual manipulation and degrades contextually appropriate behavior. We then recast prompt injection via the lens of Contextual Integrity (CI), a privacy theory that judges information flow compliance with contextual norms. This explains types of attacks that current defe
Record details
Published: 17 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Privacy · Military & security · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
US eyes ban on Chinese humanoid robots as US-China tech rivalry intensifies
SCMP Tech (HK/CN) · 23 July 2026
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
arXiv · 18 May 2026
Beyond Killer Robots: General AI Attitudes and Public Support for Military AI in Nine Countries
arXiv · 24 May 2026
Toward Agentic Governance: What Shapes LLM-Agent Intervention in Public Forums?
arXiv · 30 May 2026
Data Flow Control: Data Safety Policies for AI Agents
arXiv · 4 June 2026
End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians
arXiv · 30 April 2026
How to cite this record
ethics.ai (17 May 2026), “AI Agents May Always Fall for Prompt Injections,” evidence record 4174, https://ethics.ai/record/4174 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.