Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judg
Record details
Published: 1 July 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
arXiv · 3 July 2026
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
arXiv · 6 July 2026
Learning Action Priors for Cross-embodiment Robot Manipulation
arXiv · 24 June 2026
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
arXiv · 13 July 2026
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
arXiv · 15 July 2026
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
arXiv · 16 July 2026
How to cite this record
ethics.ai (1 July 2026), “Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity,” evidence record 345, https://ethics.ai/record/345 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.