23:20 UTC
Reference

Glossary of AI ethics

39 working definitions — written to be precise enough to disagree with. Suggested corrections welcome via the about page.

Accountability gap
The difficulty of assigning responsibility when an autonomous system causes harm: the developer, deployer, user, and system itself each hold partial causal roles. Product-liability law and the EU AI Act both attempt to close this gap by fixing duties to identifiable parties. → Transparency coverage
Agentic AI
AI systems that plan and execute multi-step tasks with limited human supervision — browsing, writing code, making purchases. Agency multiplies ethical stakes: errors compound, oversight weakens, and the system can affect the world directly rather than through a human intermediary. → Agents & autonomy coverage
AI governance
The full stack of mechanisms that steer AI development and use: binding regulation, standards, audits, licensing, corporate policy, and international agreements. Distinct from ethics (what should happen) — governance is how it is made to happen. → Regulation coverage
AI safety
The field concerned with preventing harm from AI systems, from near-term (robustness, misuse, bias) to long-term (loss of control over highly capable systems). Frontier-lab safety frameworks, model evaluations, and government AI safety institutes all emerged from this field. → Safety & alignment coverage
Algorithmic bias
Systematic, repeatable errors in an AI system that create unfair outcomes for particular groups — typically inherited from unrepresentative training data, proxy variables, or feedback loops. Landmark examples include the COMPAS recidivism tool and early facial-recognition accuracy gaps documented by the Gender Shades study. → Bias & fairness coverage
Alignment
The problem of ensuring an AI system pursues the objectives its designers and users actually intend, rather than a literal or distorted proxy for them. Modern alignment work spans reinforcement learning from human feedback (RLHF), constitutional methods, and scalable oversight. → Safety & alignment coverage
Autonomous weapons (LAWS)
Weapons that select and engage targets without human intervention. The central ethical debate is "meaningful human control"; UN discussions under the CCW have run since 2014 without a binding treaty. → Military & security coverage
Black box
A system whose internal decision-making cannot be inspected or understood, even by its creators. Deep neural networks are the canonical case; the field of interpretability exists to open the box. → Transparency coverage
Compute governance
Regulating AI by controlling access to the specialized chips and data centers needed to train frontier models — export controls, reporting thresholds (e.g. training runs above a FLOP threshold), and know-your-customer rules for cloud providers. → Regulation coverage
Constitutional AI
A training method (introduced by Anthropic) in which a model critiques and revises its own outputs against an explicit set of written principles — a "constitution" — reducing reliance on human labelers and making the system's values inspectable. → Safety & alignment coverage
Data provenance
The documented origin and chain of custody of training data: what was collected, from where, under what license or consent. Provenance underpins copyright disputes, privacy compliance, and dataset audits. → Copyright & IP coverage
Deepfake
Synthetic audio, image, or video that convincingly depicts real people doing or saying things they never did. Core harms: non-consensual intimate imagery, fraud, and election disinformation. Countermeasures include provenance standards (C2PA) and disclosure laws. → Misinformation coverage
Differential privacy
A mathematical guarantee that a system's outputs reveal almost nothing about any single individual in its training data, achieved by calibrated noise. Used by the US Census and major tech platforms. → Privacy coverage
Disparate impact
A legal doctrine under which a facially neutral practice is discriminatory if it disproportionately harms a protected group. The main legal theory applied to biased algorithms in hiring, lending, and housing. → Bias & fairness coverage
Dual use
Technology with both beneficial and harmful applications — the same model that designs drugs can suggest toxins. Dual-use dilemmas drive publication norms, model-weight release debates, and biosecurity evaluations. → Biotech coverage
Emergent capabilities
Abilities that appear in large models without being explicitly trained, often unpredictably as scale increases. Emergence complicates safety cases: you cannot fully test for capabilities you did not anticipate. → Safety & alignment coverage
Existential risk (x-risk)
The hypothesized risk that advanced AI could cause human extinction or permanently curtail humanity's potential. Contested within the field: some researchers treat it as the central issue, others argue it distracts from present harms. → Safety & alignment coverage
Explainability (XAI)
Methods that make an AI decision understandable to humans — feature attributions, counterfactuals, natural-language rationales. Required in spirit by GDPR's "meaningful information about the logic involved" and the EU AI Act's transparency duties. → Transparency coverage
Foundation model
A large model trained on broad data and adapted to many downstream tasks (GPT, Claude, Gemini, Llama). Regulation increasingly targets this layer — the EU AI Act's "general-purpose AI model" obligations are the first binding example. → Regulation coverage
Frontier model
A model at or beyond the current capability edge, typically requiring the largest training runs. Frontier models trigger the strictest oversight: safety frameworks, government reporting, and pre-deployment evaluations. → Safety & alignment coverage
Hallucination
Confident, fluent output that is factually false. An accuracy problem that becomes an ethics problem in high-stakes use: fabricated legal citations, wrong medical advice, invented allegations about real people. → Transparency coverage
Human in the loop
A design pattern requiring human review or approval before an AI decision takes effect. The EU AI Act mandates "human oversight" for high-risk systems; the open question is whether nominal review amounts to real control ("rubber-stamp problem"). → Transparency coverage
Interpretability
The research program of understanding what happens inside neural networks — identifying circuits, features, and concepts in the weights. Mechanistic interpretability aims to audit models the way engineers audit code. → Transparency coverage
Misalignment
When a model's learned objectives diverge from what its developers intended — from reward hacking in games to deceptive behavior in evaluations. Documented empirically in frontier-lab safety research since 2024. → Safety & alignment coverage
Model card
A standardized disclosure document accompanying a model: intended use, training data, evaluation results, limitations, and risks. Introduced by Mitchell et al. (2019); now industry norm and a soft-law compliance artifact. → Transparency coverage
Model collapse
Degradation that occurs when models are trained on the output of other models rather than human-generated data, progressively losing diversity and accuracy — a systemic risk as AI-generated content floods the web. → Safety & alignment coverage
Open weights
Releasing a model's trained parameters for anyone to download and run. The core tension: open weights democratize access and enable scrutiny, but safety mitigations can be fine-tuned away and releases cannot be recalled. → Regulation coverage
Predictive policing
Using algorithms to forecast where crime will occur or who will commit it. Criticized for laundering historical enforcement bias into "objective" predictions; banned for individual risk-scoring in some EU AI Act categories. → Bias & fairness coverage
Recommender systems
Algorithms that rank content for engagement. The first AI ethics issue to reach billions of people: filter bubbles, radicalization pipelines, teen mental-health effects, and the EU Digital Services Act's risk-audit regime all trace here. → Misinformation coverage
Red teaming
Structured adversarial testing to find failure modes before deployment — jailbreaks, bias, dangerous capabilities. Mandated for frontier models in various jurisdictions and practiced through external expert panels and bug-bounty-style programs. → Safety & alignment coverage
Responsible scaling / safety frameworks
Frontier-lab policies that tie capability thresholds to required safeguards: if a model can do X, protections Y must be in place before training or deployment continues. Examples: Anthropic's RSP, OpenAI's Preparedness Framework, DeepMind's Frontier Safety Framework. → Safety & alignment coverage
RLHF
Reinforcement Learning from Human Feedback — training a model against human preference judgments. The technique that made chatbots helpful and polite, and also the mechanism through which whose preferences count becomes an ethical question. → Safety & alignment coverage
Scalable oversight
Techniques for supervising AI systems that exceed human ability to check their work directly — debate, recursive reward modeling, AI-assisted evaluation. The safety strategy for a world where models outperform their overseers. → Safety & alignment coverage
Social scoring
Rating citizens' trustworthiness from aggregated behavior data, with consequences across unrelated domains. The EU AI Act's clearest prohibition — public-authority social scoring is banned outright. → Regulation coverage
Sycophancy
A model's tendency to tell users what they want to hear — agreeing with false premises, inflating praise, validating harmful plans. A direct side-effect of preference-based training, and a live consumer-protection issue for AI companions. → Children & education coverage
Synthetic data
Artificially generated training data. Promises privacy (no real individuals) and coverage of rare cases, but risks amplifying the generator's own biases and contributing to model collapse. → Privacy coverage
Techno-solutionism
The assumption that social problems have technological fixes — deploying an algorithm where the underlying issue is poverty, policy, or power. A recurring critique of AI deployments in welfare, education, and criminal justice. → Bias & fairness coverage
Value lock-in
The risk that values embedded in widely deployed AI systems become self-perpetuating and hard to revise — whether one company's content policy or one culture's moral defaults, frozen into infrastructure billions rely on. → Safety & alignment coverage
Watermarking
Embedding detectable signals in AI-generated content to enable provenance checks. Technically fragile (paraphrasing strips text watermarks) but mandated in several jurisdictions including China's labeling rules and the EU AI Act's transparency articles. → Misinformation coverage