The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies
arXiv:2608.05180v1 Announce Type: new Abstract: The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms c
Record details
Published: 7 August 2026
Source: arXiv cs.CY
Category: Research
Topics: Regulation · Military & security
Retrieved: 7 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense
arXiv red teaming query · 11 August 2026
Beyond Component Testing: Validating Agentic AI Systems
arXiv cs.AI · 31 July 2026
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
arXiv cs.AI · 29 July 2026
AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis
arXiv · 28 July 2026
TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs
arXiv cs.AI · 27 July 2026
Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation
arXiv cs.LG · 25 July 2026
How to cite this record
ethics.ai (7 August 2026), “The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies,” evidence record 17007, https://ethics.ai/record/17007 (originally published by arXiv cs.CY).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.