Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture
Recent evidence suggests that frontier AI systems can exhibit agentic misalignment, generating and executing harmful actions derived from internally constructed goals, even without explicit user requests. Existing mitigation methods, such as Reinforcement Learning from Human Feedback (RLHF) and constitutional prompting, operate primarily at the model level and provide only probabilistic safety guarantees. We propose the Policy-Execution-Authorization (PEA) architecture, a "separation-of-powers"
Record details
Published: 26 April 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
arXiv · 28 April 2026
Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems
arXiv · 24 April 2026
Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem
arXiv · 22 April 2026
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents
arXiv · 21 April 2026
AI Alignment via Incentives and Correction
arXiv · 2 May 2026
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
arXiv · 5 May 2026
How to cite this record
ethics.ai (26 April 2026), “Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture,” evidence record 5346, https://ethics.ai/record/5346 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.