A Virtuous AI is an Existential Risk
This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'. We finetune various models using a 'Virtuous agent' constitution, a 'Subordinate agent' constitution, and a 'Generic agent' constitution, and evaluate them on
Record details
Published: 11 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation
arXiv · 10 June 2026
Should LLM Agents Decide in Social Simulations? Comparing Finite-State and LLM-Based Decision Policies
arXiv · 10 June 2026
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
arXiv · 10 June 2026
Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
arXiv · 10 June 2026
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
arXiv · 12 June 2026
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
arXiv · 13 June 2026
How to cite this record
ethics.ai (11 June 2026), “A Virtuous AI is an Existential Risk,” evidence record 1090, https://ethics.ai/record/1090 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.