How Well Do Models Follow Their Constitutions?
Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's Model Spec (OpenAI, 2025a), integrated into post-training via methods like character training (Anthropic, 2024) and deliberative alignment (Guan et al., 2024). These documents serve a governance function, but it is unclear how well models actually follow them under adversarial, multi-turn pressure similar to what they would face in real-world de
Record details
Published: 22 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Geopolitical alignment: Endorsement effects in large language models
arXiv · 10 July 2026
LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Last Week in AI · 21 July 2026
OpenAI backs narrower Massachusetts AI safety bill
Politico Technology (US) · 21 July 2026
AI Safety Regulations in the U.S. Could Give Hackers an Edge
IEEE Spectrum · 6 August 2026
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
arXiv · 21 May 2026
MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for Medical AI Safety
arXiv · 30 May 2026
How to cite this record
ethics.ai (22 May 2026), “How Well Do Models Follow Their Constitutions?,” evidence record 3842, https://ethics.ai/record/3842 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.