MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing safety alignment evaluates overt requests in isolation, leaving models blind to malicious end-states that emerge from sequenced compliance with innocuous-looking requests. We introduce MOSAIC-Bench (Malicious Objectives Sequenced As Innocuous Compliance), a benchmark of 199 three-stage attack chains paired with determi
Record details
Published: 5 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents
arXiv · 7 May 2026
AI Alignment via Incentives and Correction
arXiv · 2 May 2026
Containment Verification: AI Safety Guarantees Independent of Alignment
arXiv · 9 May 2026
Guided Streaming Stochastic Interpolant Policy
arXiv · 11 May 2026
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
arXiv · 12 May 2026
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
arXiv · 28 April 2026
How to cite this record
ethics.ai (5 May 2026), “MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents,” evidence record 4953, https://ethics.ai/record/4953 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.