Evidence record 17007 · automatically gathered

The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies

arXiv:2608.05180v1 Announce Type: new Abstract: The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms c

Record details

Published: 7 August 2026
Source: arXiv cs.CY
Category: Research
Topics: Regulation · Military & security
Retrieved: 7 August 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (7 August 2026), “The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies,” evidence record 17007, https://ethics.ai/record/17007 (originally published by arXiv cs.CY).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.