Evidence record 18 · automatically gathered

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete operational detail. We propose AMT-X (Adaptive Multi-Turn Exploitation), a phase-structured multi-turn red-teaming framework. Unlike prior multi-turn attacks that rely on ad hoc escalation or free-form p

Record details

Published: 13 July 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (13 July 2026), “AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation,” evidence record 18, https://ethics.ai/record/18 (originally published by arXiv).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.