23:25 UTC
VC

Vincent Conitzer

Professor & Director, Foundations of Cooperative AI Lab

Carnegie Mellon University / University of Oxford

Research on game theory, social choice, and cooperative AI and its ethics.

Homepage / profile → Safety & alignment coverage →

Latest in the feed

From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment

arXiv:2604.04788v2 Announce Type: replace Abstract: Large language models (LLMs) could produce systematically misaligned output, from hallucinated citations to strategic deception of evaluators, yet these phenomena are studied by separate communities with incompatible terminology. We propose a unified taxonomy organized along three complementary dimensions: degree of goal-directedness (behavioral to strategic deception), object of deception, and mechanism (fabrication, omission, or pragmatic dis
arXiv cs.CY 9d ago Research Safety & alignment

Align AI to Dynamic Human-AI Workflows

Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions. In this paper, we argue for a shift from static and emulative to interactive and complementary alignment, where preferences emerge through interaction and alignment is defined not by satisfying preferences alone. We first formalize this gap by contrasting existing alignment with a
arXiv 15d ago Research Safety & alignment

Moral Decision Making Frameworks for Artificial Intelligence

The generality of decision and game theory has enabled domain-independent progress in AI research. For example, a better algorithm for finding good policies in (PO)MDPs can be instantly used in a variety of applications. But such a general theory is lacking when it comes to moral decision making. For AI applications with a moral component, are we then forced to build systems based on many ad-hoc rules? In this paper we discuss possible ways to avoid this conclusion.
OpenAlex 3455d ago Research