Research · RSS feed
New papers on fairness, safety, alignment and governance.
BiCLIP: Domain Canonicalization via Structured Geometric Transformation
Recent advances in vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities, yet adapting these models to specialized domains remains a significant challenge. Building on recent theoretical insights suggesting that independently trained VLMs are related by a canonical transformation, we extend this understanding to the concept of domains. We hypothesize that image features across disparate domains are related by a canonicalized geometric transformation that can be recove
UNBOX: Unveiling Black-box visual models with Natural-language
Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet modern vision systems are increasingly deployed as proprietary black-box APIs, exposing only output probabilities and hiding architecture, parameters, gradients, and training data. This opacity prevents meaningful auditing, bias detection, and failure analysis. Existing explanation methods assume white- or gray-box access or knowledge of the training dist
Where Do Flow Semantics Reside? A Protocol-Native Tabular Pretraining Paradigm for Encrypted Traffic Classification
Self-supervised masked modeling shows promise for encrypted traffic classification by masking and reconstructing raw bytes. Yet recent work reveals these methods fail to reduce reliance on labeled data despite costly pretraining: under frozen encoder evaluation, accuracy drops from greater than 0.9 to less than 0.47. We argue the root cause is inductive bias mismatch: flattening traffic into byte sequences destroys protocol-defined semantics. We identify three specific issues: 1) field unpredict
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
We study the implicit bias of Sharpness-Aware Minimization (SAM) when training $L$-layer linear diagonal networks on linearly separable binary classification. For linear models ($L=1$), both $\ell_\infty$- and $\ell_2$-SAM recover the $\ell_2$ max-margin classifier, matching gradient descent (GD). However, for depth $L = 2$, the behavior changes drastically -- even on a single-example dataset. For $\ell_\infty$-SAM, the limit direction depends critically on initialization and can converge to $\m
AI Misuse in Education Is a Measurement Problem: Toward a Learning Visibility Framework
The rapid integration of conversational AI systems into educational settings has intensified ethical concerns about academic integrity, fairness, and students' cognitive development. Institutional responses have largely centered on AI detection tools and restrictive policies, yet such approaches have proven unreliable and ethically contentious. This paper reframes AI misuse in education not primarily as a detection problem, but as a measurement problem rooted in the loss of visibility into the l
Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context
Large language models (LLMs) increasingly influence global digital ecosystems, yet their potential to perpetuate social and cultural biases remains poorly understood in underrepresented contexts. This study presents a systematic analysis of representational biases in seven state-of-the-art LLMs: GPT-4o-mini, Claude-3-Sonnet, Claude-4-Sonnet, Gemini-2.0-Flash, Gemini-2.0-Lite, Llama-3-70B, and Mistral-Nemo in the Nepali cultural context. Using Croissant-compliant dataset of 2400+ stereotypical an
Position: LLMs Must Use Functor-Based and RAG-Driven Bias Mitigation for Fairness
Biases in large language models (LLMs) often manifest as systematic distortions in associations between demographic attributes and professional or social roles, reinforcing harmful stereotypes across gender, ethnicity, and geography. This position paper advocates for addressing demographic and gender biases in LLMs through a dual-pronged methodology, integrating category-theoretic transformations and retrieval-augmented generation (RAG). Category theory provides a rigorous, structure-preserving
Norm-Hierarchy Transitions in Representation Learning: When and Why Neural Networks Abandon Shortcuts
Neural networks often rely on spurious shortcuts for many epochs before discovering structured representations. However, the mechanism governing when this transition occurs and whether its timing can be predicted remains unclear. Prior work shows that gradient descent converges to low norm solutions and that neural networks exhibit simplicity bias, but neither explains the timescale of the transition from shortcut features to structured representations. We introduce the Norm-Hierarchy Transition
Masking Causality and Conditional Dependence
Many regulatory and analytic problems require that a prohibited variable influence a decision only through a designated allowable channel -- a conditional-independence requirement that arises in path-specific fairness, the handling of classified information, and the regulation of trading on non-public information, among other settings. Such requirements may be enforced either stratum-by-stratum or, more commonly (and more efficiently), through a single averaged constraint on the conditional effe
Human, Algorithm, or Both? Gender Bias in Human-Augmented Recruiting
Recent years have seen rapid growth in the market for HR technology and AI-driven HR solutions in particular. This popularity has also resulted in increased attention to the negative aspects of using AI to support hiring practices, such as the risk of reinforcing existing biases against vulnerable groups based on gender or other sensitive attributes. Combining human experience with AI efficiency in making recruiting and selection decisions has the potential to help mitigate these biases, but des
Calibrated Credit Intelligence: Shift-Robust and Fair Risk Scoring with Bayesian Uncertainty and Gradient Boosting
Credit risk scoring must support high-stakes lending decisions where data distributions change over time, probability estimates must be reliable, and group-level fairness is required. While modern machine learning models improve default prediction accuracy, they often produce poorly calibrated scores under distribution shift and may create unfair outcomes when trained without explicit constraints. This paper proposes Calibrated Credit Intelligence (CCI), a deployment-oriented framework that comb
Measuring Perceptions of Fairness in AI Systems: The Effects of Infra-marginality
Differences in data distributions between demographic groups, known as the problem of infra-marginality, complicate how people evaluate fairness in machine learning models. We present a user study with 85 participants in a hypothetical medical decision-making scenario to examine two treatments: group-specific model performance and training data availability. Our results show that participants did not equate fairness with simple statistical parity. When group-specific performances were equal or u
Ambiguity Collapse by LLMs: A Taxonomy of Epistemic Risks
Large language models (LLMs) are increasingly used to make sense of ambiguous, open-textured, value-laden terms. Platforms routinely rely on LLMs for content moderation, asking them to label text based on disputed concepts like "hate speech" or "incitement"; hiring managers may use LLMs to rank who counts as "qualified"; and AI labs increasingly train models to self-regulate under constitutional-style ambiguous principles such as "biased" or "legitimate". This paper introduces ambiguity collapse
The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology
Mechanistic interpretability typically relies on post-hoc analysis of trained networks. We instead adopt an interventional approach: testing hypotheses a priori by modifying architectural topology to observe training dynamics. We study grokking - delayed generalization in Transformers trained on cyclic modular addition (Zp) - investigating if specific architectural degrees of freedom prolong the memorization phase. We identify two independent structural factors in standard Transformers: unbounde
Small Changes, Big Impact: Demographic Bias in LLM-Based Hiring Through Subtle Sociocultural Markers in Anonymised Resumes
Large Language Models (LLMs) are increasingly deployed in resume screening pipelines. Although explicit PII (e.g., names) is commonly redacted, resumes typically retain subtle sociocultural markers (languages, co-curricular activities, volunteering, hobbies) that can act as demographic proxies. We introduce a generalisable stress-test framework for hiring fairness instantiated in the Singapore context: 100 neutral job-aligned resumes are augmented into 4100 variants spanning four ethnicities and
Reimagining psychiatric care with agentic AI: promise, challenges, and a roadmap forward
Agentic artificial intelligence (AI) represents a pivotal shift in clinical decision support, moving beyond static tools by reasoning, adapting, and acting alongside clinicians. Psychiatry, grounded in subjective experience, trust, and longitudinal care, offers both an opportunity and a high-stakes testbed. Agentic systems may enhance documentation, personalize care, support continuous monitoring, and extend access, while raising risks around bias, explainability, privacy, and therapeutic allian
Governing Healthcare AI in the Real World: How Fairness, Transparency, and Human Oversight Can Coexist: A Narrative Review
Artificial intelligence (AI) is rapidly shifting from experimental pilots to mainstream clinical infrastructure, redefining how evidence, accountability, and ethics intersect in healthcare. This narrative review integrates insights from peer-reviewed studies and policy frameworks to examine seven cross-cutting aspects: bias and fairness, explainability, safety and quality, privacy and data protection, accountability and liability, human oversight, and procurement and deployment. Findings reveal
Examining human reliance on artificial intelligence in decision making
The use of Artificial Intelligence (AI) to effectively support human decision making depends on whether humans are willing to trust in, and thus rely on, AI. Understanding human reliance on AI is critical given controversial reports of AI inaccuracy and bias. Furthermore, the erroneous belief that using technology removes biases may lead to overreliance on AI. To examine humans’ reliance on AI, human participants (N = 295, Mage = 33.79) judged the authenticity of 80 faces (40 real, 40 AI-synthes
Evaluating the accuracy and reliability of AI content detectors in academic contexts
The rapid adoption of generative AI (GenAI) in higher education has intensified concerns about academic integrity, particularly for institutions serving English as a Foreign Language (EFL) learners. AI content detectors such as Turnitin and Originality are now widely used to identify potential misuse of GenAI in student writing, yet their accuracy, consistency, and fairness remain to be proven. This study evaluates the reliability of these two commercial detectors using a balanced dataset of 192
Governing the blue economy in arid coastal regions: opportunities, constraints, and stakeholder perspectives from the Eastern Province coast of Saudi Arabia
Introduction The blue economy has emerged as a strategic framework for aligning marine-based economic development with environmental sustainability and social equity. Empirical evidence from arid and industrialized coastal regions, however, remains limited. Methods This study employs a convergent mixed-methods design using a structured questionnaire administered to 404 stakeholders across the Eastern Province coastline of Saudi Arabia, complemented by qualitative open-ended responses. Quantitati
Total cholesterol, high-density lipoprotein, and glucose (CHG) index and diabetic retinopathy in middle-aged and elderly Chinese adults with diabetes: a cross-sectional study
Objective: Evidence regarding the association between the total cholesterol, high-density lipoprotein, and glucose (CHG) index and diabetic retinopathy (DR) remains limited. This study aimed to explore the relationship between CHG and the prevalence of DR and evaluate its discriminative ability for DR. Methods: This cross-sectional study analyzed data from 1,909 individuals with diabetes mellitus (DM), aged 45-90 years, whose information was collected between August and December 2011. To determi
Six Institutional Intervention Areas to Support Ethical and Effective Student Use of Generative AI in Higher Education: A Narrative Review
The integration of generative AI tools, such as ChatGPT, Gemini, and DeepSeek, into higher education offers transformative opportunities for personalised learning and academic productivity. However, their unregulated use raises concerns about academic integrity, critical thinking, and educational equity. This systematic review synthesises insights from 96 peer-reviewed articles, identifying six key intervention themes, namely, curriculum integration, policy and governance, faculty development, s
Intersectional biases in narratives produced by open-ended prompting of generative language models
The rapid deployment of generative language models has raised concerns about social biases affecting the well-being of diverse consumers. The extant literature on generative language models has primarily examined bias via explicit identity prompting. However, prior research on bias in language-based technology platforms has shown that discrimination can occur even when identity terms are not specified explicitly. Here, we advance studies of generative language model bias by considering a broader
Transforming clinical reasoning—the role of AI in supporting human cognitive limitations
Clinical reasoning is foundational to medical practice, requiring clinicians to synthesise complex information, recognise patterns, and apply causal reasoning to reach accurate diagnoses and guide patient management. However, human cognition is inherently limited by factors such as limitations in working memory capacity, constraints in cognitive load, a general reliance on heuristics; with an inherent vulnerability to biases including anchoring, availability bias, and premature closure. Cognitiv
Differential privacy for medical deep learning: methods, tradeoffs, and deployment implications
Differential privacy (DP) is a prominent technique for protecting sensitive patient data in medical deep learning (DL), yet deploying it without compromising clinical utility or equity remains challenging. This scoping review synthesizes applications of DP in medical DL across centralized and federated settings. A structured search identified 74 eligible studies published through March 2025. Across modalities and tasks, DP, especially via DP-SGD, can maintain clinically acceptable performance un
Discrimination, artificial intelligence, and algorithmic decision-making
Artificial intelligence (AI) has a huge impact on our personal lives and also on our democratic society as a whole. While AI offers vast opportunities for the benefit of people, its potential to embed and perpetuate bias and discrimination remains one of the most pressing challenges deriving from its increasing use. This new study, which was prepared by Prof. Frederik Zuiderveen Borgesius for the Anti-discrimination Department of the Council of Europe, elaborates on the risks of discrimination c
Exploring automation bias in human–AI collaboration: a review and implications for explainable AI
Abstract As Artificial Intelligence (AI) becomes increasingly embedded in high-stakes domains such as healthcare, law, and public administration, automation bias (AB)—the tendency to over-rely on automated recommendations—has emerged as a critical challenge in human–AI collaboration. While previous reviews have examined AB in traditional computer-assisted decision-making, research on its implications in modern AI-driven work environments remains limited. To address this gap, this research system
Ethical and regulatory challenges of Generative AI in education: a systematic review
Introduction Generative Artificial Intelligence (GenAI) is transforming education by enabling personalized learning and more efficient teaching practices. However, it raises critical ethical concerns, including data privacy, algorithmic bias, and educational inequality, requiring comprehensive regulatory frameworks and pedagogical strategies. Methods A Systematic Literature Review (SLR) was conducted, analyzing 53 peer-reviewed articles published between 2020 and 2024. The search was performed i
Transparency in the Reporting of Artificial Intelligence – The TITAN Guideline
The use of AI in research and the literature is increasing. The need for transparency is clear. Here we present a guideline to transparently report the use of AI in any manuscript in general. The guideline items cover; declaration, purpose and scope, AI tools and configuration, data inputs and safeguards, human oversight and verification, bias, ethics and regulatory compliance and reproducibility and transparency. These items have been confirmed in a recent Delphi consensus exercise with high pa
Comparing Generative AI and teacher feedback: student perceptions of usefulness and trustworthiness
The rapid integration of Generative Artificial Intelligence (GenAI) into educational contexts has presented both opportunities and challenges for students seeking and using feedback. While AI-generated feedback can offer increased access, timely responses and personalised insights, concerns about the quality of AI-generated feedback still persist, including issues of bias, factual inaccuracies, and homogenisation. This study investigates how students use, value and trust AI-generated feedback co
Integrating Digital Health Innovations to Achieve Universal Health Coverage: Promoting Health Outcomes and Quality Through Global Public Health Equity
Digital health innovations are reshaping global healthcare systems by enhancing access, efficiency, and quality of care. Technologies such as artificial intelligence, telemedicine, mobile health applications, and big data analytics have been widely applied to support disease surveillance, enable remote care, and improve clinical decision making. This review critically identifies persistent implementation challenges that hinder the equitable adoption of digital health solutions, such as the digit
CONSORT 2025 explanation and elaboration: updated guideline for reporting randomised trials
This is comment on: Hopewell S, et al. CONSORT 2025 explanation and elaboration: updated guideline for reporting randomised trials. BMJ. 2025 Apr 14;389:e081124. https://pubmed.ncbi.nlm.nih.gov/40228832 The CONSORT (2025) statement [1] writes about blinding as follows: “Unblinded outcome assessors may differentially assess subjective outcomes, and unblinded data analysts may introduce bias through the choice of analytical strategies, such as the selection of favourable time points or outcomes an
Generalization bias in large language model summarization of scientific research
Artificial intelligence chatbots driven by large language models (LLMs) have the potential to increase public science literacy and support scientific research, as they can quickly summarize complex scientific information in accessible terms. However, when summarizing scientific texts, LLMs may omit details that limit the scope of research conclusions, leading to generalizations of results broader than warranted by the original study. We tested 10 prominent LLMs, including ChatGPT-4o, ChatGPT-4.5
Shaping the Future of Healthcare: Ethical Clinical Challenges and Pathways to Trustworthy AI
Background/Objectives: Artificial intelligence (AI) is transforming healthcare, enabling advances in diagnostics, treatment optimization, and patient care. Yet, its integration raises ethical, regulatory, and societal challenges. Key concerns include data privacy risks, algorithmic bias, and regulatory gaps that struggle to keep pace with AI advancements. This study aims to synthesize a multidisciplinary framework for trustworthy AI in healthcare, focusing on transparency, accountability, fairne
AI Ethics: Integrating Transparency, Fairness, and Privacy in AI Development
The expansion of Artificial Intelligence in sectors such as healthcare, finance, and communication has raised critical ethical concerns surrounding transparency, fairness, and privacy. Addressing these issues is essential for the responsible development and deployment of AI systems. This research establishes a comprehensive ethical framework that mitigates biases and promotes accountability in AI technologies. A comparative analysis of international AI policy frameworks from regions including th
Multi-omics approaches for understanding gene-environment interactions in noncommunicable diseases: techniques, translation, and equity issues
Non-communicable diseases (NCDs) such as cardiovascular diseases, chronic respiratory diseases, cancers, diabetes, and mental health disorders pose a significant global health challenge, accounting for the majority of fatalities and disability-adjusted life years worldwide. These diseases arise from the complex interactions between genetic, behavioral, and environmental factors, necessitating a thorough understanding of these dynamics to identify effective diagnostic strategies and interventions
Retrieval-augmented generation for generative artificial intelligence in health care
Abstract Generative artificial intelligence has brought disruptive innovations in health care but faces certain challenges. Retrieval-augmented generation (RAG) enables models to generate more reliable content by leveraging the retrieval of external knowledge. In this perspective, we analyze the possible contributions that RAG could bring to health care in equity, reliability, and personalization. Additionally, we discuss the current limitations and challenges of implementing RAG in medical scen
Generative AI in Higher Education: Balancing Innovation and Integrity
Generative Artificial Intelligence (GenAI) is rapidly transforming the landscape of higher education, offering novel opportunities for personalised learning and innovative assessment methods. This paper explores the dual-edged nature of GenAI's integration into educational practices, focusing on both its potential to enhance student engagement and learning outcomes and the significant challenges it poses to academic integrity and equity. Through a comprehensive review of current literature, we e
How human–AI feedback loops alter human perceptual, emotional and social judgements
Artificial intelligence (AI) technologies are rapidly advancing, enhancing human capabilities across various fields spanning from finance to medicine. Despite their numerous advantages, AI systems can exhibit biased judgements in domains ranging from perception to emotion. Here, in a series of experiments (n = 1,401 participants), we reveal a feedback loop where human-AI interactions alter processes underlying human perceptual, emotional and social judgements, subsequently amplifying biases in h