Evaluating Large Language Models in a Complex Hidden Role Game
Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments. This work investigates the reasoning, persuasion, and deceptive capabilities of LLMs within the social deduction game Secret Hitler. I introduce an open-source framework and novel metrics to measure performance: Role Identification Accuracy, Deception Retention Rate, and Game State Impact Rate. By benchmarking models against rule-based algorithms a
Record details
Published: 9 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Environment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
arXiv · 6 May 2026
Amplifying the storm: Climate disinformation dynamics during natural disasters on right-wing extremist Telegram channels
HKS Misinformation Review · 23 July 2026
Climate alignment of finance: different policy playbooks and untapped investment opportunities
OECD · 5 July 2026
Microsoft built an agentic security system with red, blue, and green team AI agents. It enters public preview August 3.
The Next Web AI · 27 July 2026
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
arXiv · 9 April 2026
From Safety Risk to Design Principle: Peer-Preservation in Multi-Agent LLM Systems and Its Implications for Orchestrated Democratic Discourse Analysis
arXiv · 9 April 2026
How to cite this record
ethics.ai (9 April 2026), “Evaluating Large Language Models in a Complex Hidden Role Game,” evidence record 6123, https://ethics.ai/record/6123 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.