NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms
Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as operators of a safety-critical system, instantiated in a simulated nuclear power plant control room. A five-role operator team, each backed by a configurable LLM, runs a plant governed by six critical s
Record details
Published: 18 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
PrivacyAlign: Contextual Privacy Alignment for LLM Agents
arXiv · 19 June 2026
RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills
arXiv · 16 June 2026
Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion
arXiv · 20 June 2026
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
arXiv · 15 June 2026
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
arXiv · 15 June 2026
Agentic Framework for Deep Learning workload migration via In-Context Learning
arXiv · 14 June 2026
How to cite this record
ethics.ai (18 June 2026), “NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms,” evidence record 797, https://ethics.ai/record/797 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.