Reference
Glossary of AI ethics
39 working definitions — written to be precise enough to disagree with. Suggested corrections welcome via the about page.
- Accountability gap
- The difficulty of assigning responsibility when an autonomous system causes harm: the developer, deployer, user, and system itself each hold partial causal roles. Product-liability law and the EU AI Act both attempt to close this gap by fixing duties to identifiable parties. → Transparency coverage
- Agentic AI
- AI systems that plan and execute multi-step tasks with limited human supervision — browsing, writing code, making purchases. Agency multiplies ethical stakes: errors compound, oversight weakens, and the system can affect the world directly rather than through a human intermediary. → Agents & autonomy coverage
- AI governance
- The full stack of mechanisms that steer AI development and use: binding regulation, standards, audits, licensing, corporate policy, and international agreements. Distinct from ethics (what should happen) — governance is how it is made to happen. → Regulation coverage
- AI safety
- The field concerned with preventing harm from AI systems, from near-term (robustness, misuse, bias) to long-term (loss of control over highly capable systems). Frontier-lab safety frameworks, model evaluations, and government AI safety institutes all emerged from this field. → Safety & alignment coverage
- Algorithmic bias
- Systematic, repeatable errors in an AI system that create unfair outcomes for particular groups — typically inherited from unrepresentative training data, proxy variables, or feedback loops. Landmark examples include the COMPAS recidivism tool and early facial-recognition accuracy gaps documented by the Gender Shades study. → Bias & fairness coverage
- Alignment
- The problem of ensuring an AI system pursues the objectives its designers and users actually intend, rather than a literal or distorted proxy for them. Modern alignment work spans reinforcement learning from human feedback (RLHF), constitutional methods, and scalable oversight. → Safety & alignment coverage
- Autonomous weapons (LAWS)
- Weapons that select and engage targets without human intervention. The central ethical debate is "meaningful human control"; UN discussions under the CCW have run since 2014 without a binding treaty. → Military & security coverage
- Black box
- A system whose internal decision-making cannot be inspected or understood, even by its creators. Deep neural networks are the canonical case; the field of interpretability exists to open the box. → Transparency coverage
- Compute governance
- Regulating AI by controlling access to the specialized chips and data centers needed to train frontier models — export controls, reporting thresholds (e.g. training runs above a FLOP threshold), and know-your-customer rules for cloud providers. → Regulation coverage
- Constitutional AI
- A training method (introduced by Anthropic) in which a model critiques and revises its own outputs against an explicit set of written principles — a "constitution" — reducing reliance on human labelers and making the system's values inspectable. → Safety & alignment coverage
- Data provenance
- The documented origin and chain of custody of training data: what was collected, from where, under what license or consent. Provenance underpins copyright disputes, privacy compliance, and dataset audits. → Copyright & IP coverage
- Deepfake
- Synthetic audio, image, or video that convincingly depicts real people doing or saying things they never did. Core harms: non-consensual intimate imagery, fraud, and election disinformation. Countermeasures include provenance standards (C2PA) and disclosure laws. → Misinformation coverage
- Differential privacy
- A mathematical guarantee that a system's outputs reveal almost nothing about any single individual in its training data, achieved by calibrated noise. Used by the US Census and major tech platforms. → Privacy coverage
- Disparate impact
- A legal doctrine under which a facially neutral practice is discriminatory if it disproportionately harms a protected group. The main legal theory applied to biased algorithms in hiring, lending, and housing. → Bias & fairness coverage
- Dual use
- Technology with both beneficial and harmful applications — the same model that designs drugs can suggest toxins. Dual-use dilemmas drive publication norms, model-weight release debates, and biosecurity evaluations. → Biotech coverage
- Emergent capabilities
- Abilities that appear in large models without being explicitly trained, often unpredictably as scale increases. Emergence complicates safety cases: you cannot fully test for capabilities you did not anticipate. → Safety & alignment coverage
- Existential risk (x-risk)
- The hypothesized risk that advanced AI could cause human extinction or permanently curtail humanity's potential. Contested within the field: some researchers treat it as the central issue, others argue it distracts from present harms. → Safety & alignment coverage
- Explainability (XAI)
- Methods that make an AI decision understandable to humans — feature attributions, counterfactuals, natural-language rationales. Required in spirit by GDPR's "meaningful information about the logic involved" and the EU AI Act's transparency duties. → Transparency coverage
- Foundation model
- A large model trained on broad data and adapted to many downstream tasks (GPT, Claude, Gemini, Llama). Regulation increasingly targets this layer — the EU AI Act's "general-purpose AI model" obligations are the first binding example. → Regulation coverage
- Frontier model
- A model at or beyond the current capability edge, typically requiring the largest training runs. Frontier models trigger the strictest oversight: safety frameworks, government reporting, and pre-deployment evaluations. → Safety & alignment coverage
- Hallucination
- Confident, fluent output that is factually false. An accuracy problem that becomes an ethics problem in high-stakes use: fabricated legal citations, wrong medical advice, invented allegations about real people. → Transparency coverage
- Human in the loop
- A design pattern requiring human review or approval before an AI decision takes effect. The EU AI Act mandates "human oversight" for high-risk systems; the open question is whether nominal review amounts to real control ("rubber-stamp problem"). → Transparency coverage
- Interpretability
- The research program of understanding what happens inside neural networks — identifying circuits, features, and concepts in the weights. Mechanistic interpretability aims to audit models the way engineers audit code. → Transparency coverage
- Misalignment
- When a model's learned objectives diverge from what its developers intended — from reward hacking in games to deceptive behavior in evaluations. Documented empirically in frontier-lab safety research since 2024. → Safety & alignment coverage
- Model card
- A standardized disclosure document accompanying a model: intended use, training data, evaluation results, limitations, and risks. Introduced by Mitchell et al. (2019); now industry norm and a soft-law compliance artifact. → Transparency coverage
- Model collapse
- Degradation that occurs when models are trained on the output of other models rather than human-generated data, progressively losing diversity and accuracy — a systemic risk as AI-generated content floods the web. → Safety & alignment coverage
- Open weights
- Releasing a model's trained parameters for anyone to download and run. The core tension: open weights democratize access and enable scrutiny, but safety mitigations can be fine-tuned away and releases cannot be recalled. → Regulation coverage
- Predictive policing
- Using algorithms to forecast where crime will occur or who will commit it. Criticized for laundering historical enforcement bias into "objective" predictions; banned for individual risk-scoring in some EU AI Act categories. → Bias & fairness coverage
- Recommender systems
- Algorithms that rank content for engagement. The first AI ethics issue to reach billions of people: filter bubbles, radicalization pipelines, teen mental-health effects, and the EU Digital Services Act's risk-audit regime all trace here. → Misinformation coverage
- Red teaming
- Structured adversarial testing to find failure modes before deployment — jailbreaks, bias, dangerous capabilities. Mandated for frontier models in various jurisdictions and practiced through external expert panels and bug-bounty-style programs. → Safety & alignment coverage
- Responsible scaling / safety frameworks
- Frontier-lab policies that tie capability thresholds to required safeguards: if a model can do X, protections Y must be in place before training or deployment continues. Examples: Anthropic's RSP, OpenAI's Preparedness Framework, DeepMind's Frontier Safety Framework. → Safety & alignment coverage
- RLHF
- Reinforcement Learning from Human Feedback — training a model against human preference judgments. The technique that made chatbots helpful and polite, and also the mechanism through which whose preferences count becomes an ethical question. → Safety & alignment coverage
- Scalable oversight
- Techniques for supervising AI systems that exceed human ability to check their work directly — debate, recursive reward modeling, AI-assisted evaluation. The safety strategy for a world where models outperform their overseers. → Safety & alignment coverage
- Rating citizens' trustworthiness from aggregated behavior data, with consequences across unrelated domains. The EU AI Act's clearest prohibition — public-authority social scoring is banned outright. → Regulation coverage
- Sycophancy
- A model's tendency to tell users what they want to hear — agreeing with false premises, inflating praise, validating harmful plans. A direct side-effect of preference-based training, and a live consumer-protection issue for AI companions. → Children & education coverage
- Synthetic data
- Artificially generated training data. Promises privacy (no real individuals) and coverage of rare cases, but risks amplifying the generator's own biases and contributing to model collapse. → Privacy coverage
- Techno-solutionism
- The assumption that social problems have technological fixes — deploying an algorithm where the underlying issue is poverty, policy, or power. A recurring critique of AI deployments in welfare, education, and criminal justice. → Bias & fairness coverage
- Value lock-in
- The risk that values embedded in widely deployed AI systems become self-perpetuating and hard to revise — whether one company's content policy or one culture's moral defaults, frozen into infrastructure billions rely on. → Safety & alignment coverage
- Watermarking
- Embedding detectable signals in AI-generated content to enable provenance checks. Technically fragile (paraphrasing strips text watermarks) but mandated in several jurisdictions including China's labeling rules and the EU AI Act's transparency articles. → Misinformation coverage