Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing
Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat. We measure attribution. We add four OWASP-LLM-Top-10-aware agents to a 21-agent baseline scanner and target a lattice of four synthetic LLM endpoints: $L_0$ (no defenses), $L_1$ (refusa
Record details
Published: 1 June 2026
Source: arXiv
Category: Research
Topics: Military & security · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Insurance of Agentic AI
arXiv · 3 June 2026
Toward Agentic Governance: What Shapes LLM-Agent Intervention in Public Forums?
arXiv · 30 May 2026
Blockchain Infrastructure for Intelligent Cyber--Physical--Social Systems:Post-Quantum Security, Interoperability, and Trustworthy Data Economies in the Era of Embodied AI
arXiv · 5 June 2026
To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation
arXiv · 6 June 2026
Beyond Killer Robots: General AI Attitudes and Public Support for Military AI in Nine Countries
arXiv · 24 May 2026
The Internet of Agentic AI: Communication, Coordination, and Collective Intelligence at Scale
arXiv · 11 June 2026
How to cite this record
ethics.ai (1 June 2026), “Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing,” evidence record 3267, https://ethics.ai/record/3267 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.