Research · RSS feed
New papers on fairness, safety, alignment and governance.
Logic, Optimization, and Artificial Intelligence
Logic and optimization can, in combination, make valuable contributions to rule-based AI. Logic is the obvious medium for encoding a rule base and drawing inferences from it, while optimization provides a powerful technology for computing inferences. Their combination has taken on new relevance amid a growing concern for transparency in AI. which is important for reproducibility, explainability, trustworthiness, and fairness. Rule-based AI provides a natural solution to transparency that is beco
Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies
Online encyclopedias shape political opinion and, through it, democratic discourse. In late 2025, Grokipedia was released, an encyclopedia written entirely by the LLM Grok. One motivation behind the project was to provide an unbiased alternative to Wikipedia, which has faced accusations of "left-wing" and "liberal" bias. But does an encyclopedia written by an LLM deliver greater neutrality, or does it simply embed a different ideology? We conduct a large-scale political bias study on Grokipedia
Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers
Per-subgroup fairness audits of medical image classifiers face a sample-size problem: minority subgroups in held-out test sets have so few samples that the resulting confidence intervals on per-subgroup performance are wider than the bias the audit is meant to detect. We argue that a demographically-conditioned synthetic generator can do both: mitigate bias on the training side and detect bias on the evaluation side. Working on COVID-19 chest CT classification with an end-to-end fine-tuned Stabl
Latent Trajectory Discrimination for AI-Generated Text Detection
Most existing approaches to AI-Generated Text Detection (AIGTD) treat documents as static objects and base their decisions on aggregate statistics or globally compressed embeddings. However, this perspective overlooks the inherently dynamic nature of autoregressive generation, where content evolves progressively through the latent space. In this paper, we reformulate AIGTD as the problem of distinguishing between latent generation trajectories. Instead of relying on static representations, we mo
The Misclassification of Autistic Writing as AI-Generated
Recent findings suggest that detection models for artificial intelligence (AI) cannot accurately identify AI-generated text and may exhibit bias against certain minority groups. In the present study, anecdotal claims that autistic writers more often have their work flagged as AI-generated are examined empirically. A corpus of approximately 60,000 Reddit posts split into "likely-autistic" and "general-Reddit" subcorpora is used to compare the distribution of probabilities output by the OpenAI GPT
Grad2Fair: A Gradient-driven Approach for Graph Fairness without Demographics
Graph neural networks (GNNs) frequently encounter group fairness issues, often yielding biased predictions against specific demographic groups defined by sensitive attributes such as gender or race. While this challenge has motivated extensive research, most existing solutions rely on the strong assumption that demographics are fully available. To bypass this strict requirement, a few recent studies have attempted to use predicted demographics as proxies to enforce fairness constraints. However,
Auditing Fairness-Privacy Trade-offs: Subpopulation-Level Effects of Fairness-Enhancing Algorithms
Machine learning (ML) models deployed in sensitive domains such as healthcare, law enforcement, and finance must satisfy not only utility requirements but also fairness and privacy guarantees. While prior work has largely examined how privacy-preserving techniques affect fairness, the inverse question-how fairness-enhancing algorithms influence privacy leakage-remains underexplored. We present the first comprehensive study of how fairness interventions affect membership inference privacy risks a
Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays
This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in "AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models" (Gayed, 2026), which was fine-tuned on 480 argumentative essays from two prompts, we evaluate scoring accuracy on the full TOEFL11
When Social Media Memes Become Mean to Peers: Discrimination Recognition and Group Norms in Adolescent Bullying
Publication date: Available online 15 July 2026 Source: Computers in Human Behavior Author(s): Rongyi Chen, Qing Xiao, Shike Lin, Jingjia Xiao, Menghan Yin, Yucong Ma, Bingbing Zhang, Hua Zhong
AI Alignment Amplifies the Role of Race, Gender, and Disability in Hiring Decisions
arXiv:2605.13866v2 Announce Type: replace Abstract: Humans increasingly delegate consequential decisions to language models, yet whether these systems reproduce or reshape human patterns of discrimination remains unclear. Here, across 29 models and 177 occupations covering nearly half of U.S. employment, we show that language models incorporate demographics into hiring decisions, advantaging female and Black candidates while penalising disabled candidates, with effect sizes comparable to six mon
Designing AI-resilient assessment in higher education: a four-pillar conceptual framework
Generative AI tools can produce polished academic text on demand, undermining the validity of assessments that treat written submissions as evidence of individual learning. Detection-based countermeasures have demonstrated variable accuracy and equity concerns. This paper does not report empirical outcomes or validation data. It proposes a framework for AI-resilient assessment that shifts evaluation from product quality to demonstrable reasoning, decision-making, and ownership of learning. The f
Developing a Core Outcome Set for the Evaluation of Remote Patient Monitoring Interventions Using the Sextuple Aim: Modified Delphi Study
Background: The rapid expansion of remote patient monitoring (RPM) interventions highlights the need for their comparison and evaluation. Current evaluation frameworks often fail to capture a multistakeholder perspective. Traditional health technology assessment approaches emphasize health and economic outcomes, providing an incomplete picture of RPM’s broader impact. The Sextuple Aim, encompassing health outcomes, costs, patient and provider experience, equity, and sustainability, offers a more
Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)
As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, and educational equity. Existing measures of reliance were developed inductively, focused on discrete problem-solving tasks, and validated mainly with homogeneous samples. This study developed and validated the GenAI Reliance Types Scale (GenAI-RTS), a 20-item instrumen
NodeImport: Imbalanced Node Classification with Node Importance Assessment
In real-world applications, node classification on graphs often faces the challenge of class imbalance, where majority classes dominate training, resulting in biased model performance. Traditional GNNs often struggle in such scenarios, as they tend to overfit to majority classes while underrepresenting minority classes. Existing solutions, which either prioritize nodes based on class size or synthesize new nodes for minority classes, often fall short of effectively addressing this imbalance issu
Implementations of Quantum and Classical Topology-Aligned Architectures for Molecular Property Prediction
For low-data and resource-constrained regimes typical of quantum chemistry, parameter-efficient learning is a key objective. Here, we propose a topology-aligned inductive bias in which the model architecture mirrors the molecular bond graph: atoms map to a fixed register of computational units, and bonds determine which pairs interact through shared learnable parameters. This principle is instantiated in two architectures: a variational quantum circuit (Iso-QGNN), and a parameter-matched classic
Price of Fairness in Bandits: A Tight Minimax Characterization
In bandit problems, standard regret-minimizing algorithms treat exploration as an amortized cost, which can expose early participants to unfair ex-ante losses in settings such as clinical trials. Recent work addresses this by evaluating the sequence of per-round expected rewards through the generalized $p$-mean, interpolating between utilitarian welfare ($p=1$), Nash welfare ($p\to0$), and Rawlsian fairness ($p\to-\infty$). Although tight guarantees are known for $p\ge0$, the strictly fair regim
All too perfect: bias and aspiration in persona generation with LLMs
Synthetic data generated by large language models plays a central role in the training and alignment process of other AI systems. However, this process also risks inheriting the structural biases of organic corpora and embedding new biases that stem from the design choices underlying the data creation process. This paper examines the systematic biases that emerge when large language models (LLMs) are tasked with generating synthetic personas. We introduce a reproducible, minimally conditioned pi
Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference
Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (delimiters or whitespace) carries an order of magnitude more energy than any content role, and structural KEY to
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models
The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models. In this paper, we use a crowdsourced study to show that providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model, even when the model is patently discriminatory and unfair. Our results suggest that XAI designers and developers n
Exploring the influence of missing data imputation in group fairness metrics
Publication date: August 2026 Source: Artificial Intelligence, Volume 357 Author(s): Arthur Dantas Mangussi, Ricardo Cardoso Pereira, Miriam Seoane Santos, Ana Carolina Lorena, Mykola Pechenizkiy, Pedro Henriques Abreu
Substratism: Conceptualizing and measuring moral bias against AI
Publication date: December 2026 Source: Computers in Human Behavior, Volume 185 Author(s): Ali Ladak, Janet V.T. Pauketat, Jacy Reese Anthis, Steve Loughnan, Matti Wilks
Digital transformation and social equity: Bridging the gap in post-pandemic societies
Publication date: August 2026 Source: Technology in Society, Volume 87 Author(s): Lifeng Lai, Ted Brader
Transparency, neutrality, voice, and respect: How procedural fairness considerations affect AI acceptability in algorithmic societies
Publication date: August 2026 Source: Technology in Society, Volume 87 Author(s): Pedro C. Magalhães, Sveinung Arnesen, Christoph Kern, Pascal D. Koenig, Daniel S. Schiff, Tom R. Tyler
Attentional Priority for Social Reward is Modulated by Attentional Bias Toward Game
Publication date: Available online 10 July 2026 Source: Computers in Human Behavior Author(s): Dongyu Liu, Xinyu Zhang, Boxiang Li, Christian Montag, Jon D. Elhai, Haibo Yang
Enhancing fairness and transparency in student project evaluation: A spherical fuzzy alternative prioritization and assessment system-based decision support
Publication date: August 2026 Source: Technology in Society, Volume 87 Author(s): Hamide Özyürek, Karahan Kara, Galip Cihan Yalçın, Zeynep Baysal, Ufuk Türen, Vladimir Simic, Mustafa Polat, Dragan Pamucar
Exploring attribution bias in LLMs: Social influences and prompt-based mitigation
Publication date: August 2026 Source: Technology in Society, Volume 87 Author(s): Leilei Jiang, Jie Cao, Guixiang Zhu, Yuyao Wang
The impact of item-writing flaws on difficulty and discrimination in item response theory
Publication date: December 2026 Source: Computers and Education: Artificial Intelligence, Volume 11 Author(s): Robin Schmucker, Steven Moore
Beyond “painting in pink”: A critical case study of all-girls generative AI workshops in a European makerspace and implications for gender equity in computing
Publication date: June 2026 Source: Computers and Education: Artificial Intelligence, Volume 10 Author(s): Qian Liu, Louise Archer, Meghna Nag Chowdhuri, Esme Freedman, Jennifer DeWitt
Beyond binary outcomes: Evaluating and mitigating bias in national standardized test score prediction
Publication date: June 2026 Source: Computers and Education: Artificial Intelligence, Volume 10 Author(s): Lin Li, Namrata Srivatava, Jia Rong, Quanlong Guan, Dragan Gašević, Guanliang Chen
Bias and representation in AI generated text-to-image in education: A systematic review
Publication date: June 2026 Source: Computers and Education: Artificial Intelligence, Volume 10 Author(s): Lilach Alon, Dorit Hadar Shoval, Inbar Levkovich
Toward Localizing and Repairing Bias in Transformer Attention Heads
Transformer language models are increasingly used as software components, yet biased outputs remain difficult to localize and repair inside the model. Existing fairness testing and repair methods largely operate at the input-output or retraining level, while recent work suggests that bias-related behavior can concentrate in a small set of attention heads. This paper studies whether attention heads can be localized and repaired through a targeted inference-time intervention. We introduce ROBIN, a
HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition
This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition. For multi-task learning with simultaneous prediction of valence, arousal, facial expressions, and action units on s-Aff-Wild2 dataset, we use frozen lightweight facial extractors, MT-EmotiDDAMFN and MT-EmotiEffNet-B0, with separate heads and systematic post-processing: temporal Gaussian smoothing, per-class expression bias, AffectNet blending, per-AU threshold tuning, and weighted backbone
Accuracy and Normalized Accuracy under Length Bias: Analysis, Guidelines, and a Bayesian Alternative
Multiple-choice benchmarks that rank candidate completions by conditional log-probability suffer from a length bias: because log-probabilities sum over tokens, longer answers tend to be penalized relative to shorter ones in practice. A common mitigation is to normalize scores by completion length, but we show empirically that this heuristic frequently over-corrects, introducing a bias toward longer answers instead. We first analyze these scoring rules, characterizing when standard and length-nor
Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?
As Large Language Models (LLMs) are increasingly deployed as autonomous agents in high-stakes domains, understanding contextual factors that may modulate their decision-making becomes critical. While LLMs are trained to perceive and resonate with users' emotions, it remains unclear whether induced emotion can influence their sequential decision-making. We investigate this question using the Iowa Gambling Task (IGT), a classic psychological paradigm for studying decision-making under uncertainty,
DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs
arXiv:2607.11228v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities. We introduce DeepBias, an adaptive framework for the in-depth probing of social b
Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis
arXiv:2607.11314v1 Announce Type: new Abstract: The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science education. In most systems, a general-track subject -- Digital Literacy, ICT, TIC, or SNT -- bears the weight of universal AI literacy, while a specialist Informatics course serves STEM pathways separately. Yet the content and depth of the general track are shaped by governance decisions made largely with reference to
Neutralizing Structural Inequality in the Nigerian FinTech Sector
arXiv:2607.10317v1 Announce Type: cross Abstract: Algorithmic decision systems in financial services often rely on data proxies that inadvertently encode structural inequalities. This paper introduces a hierarchical human-AI triage model for Point of Sale fraud detection in the Nigerian FinTech sector. Adopting a We Are All Equal worldview, we address the challenge of discrimination laundering, wherein the system misinterprets infrastructure related aleatoric noise such as rural network timeouts
How Data Narratives Go Wrong: A Taxonomy of Issues Across the Data Communication Process
arXiv:2607.10523v1 Announce Type: cross Abstract: Data narratives increasingly shape public understanding, but their failures are rarely just isolated factual errors or deceptive charts. Instead, they emerge through a broader meaning-making process in which quantitative evidence is transformed into claims, representations, and arguments. While prior work has examined these failures across disparate fields (e.g., statistics, visualization, and fact-checking), the community lacks a holistic lens t
The Benchmark Ceiling: Human Judgment, Evaluation Scarcity, and the Political Economy of AI Capability Measurement
arXiv:2607.01254v2 Announce Type: replace Abstract: Benchmarks are the primary instruments through which AI capability is measured, compared, and governed. This paper argues that the validity of frontier AI benchmarks is a function of the quality of human judgment embedded in their construction, and that this quality is structurally scarce in ways that standard scaling narratives obscure. As foundation models approach ceiling performance on existing evaluation suites, discriminating signal conce
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
arXiv:2606.28186v2 Announce Type: replace-cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual representations, providing limited evidence about the cognitive processes that make items difficult. We argue that difficulty should be viewed not only as a property of item text, but also as an observable consequenc