Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this work, we identify a failure mode, deductive stereotyping, in which models apply population-level statistical regularities to individual cases, producing logically coherent yet socially biased inferences. We provide a statistical interpretation of this phenomenon. To steer models toward fairness-aware reasoning, we propo
Record details
Published: 30 June 2026
Source: arXiv
Category: Research
Topics: Bias & fairness
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs
arXiv · 30 June 2026
Test-Time Verification for Text-to-SQL via Outcome Reward Models
arXiv · 29 June 2026
AETDICE: Unified Framework and Offline Optimization for Nonlinear Multi-Objective RL
arXiv · 30 June 2026
Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models
arXiv · 30 June 2026
Sequential Fairness Auditing with Limited Output Access
arXiv · 29 June 2026
Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification
arXiv · 29 June 2026
How to cite this record
ethics.ai (30 June 2026), “Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG,” evidence record 426, https://ethics.ai/record/426 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.