Topic · updated daily · RSS feed for this topic
Children & education
AI in schools, chatbots and minors, companion-app harms and edtech ethics — tracked daily.
GIGO and the Human Processor: A CHO Review
This document is a peer review by Go Kian Tik — GKT (Mbah Hogi Bejo™) of Meriel B.’s 10-day research series. Meriel B. is an AI Ethics Researcher, AI Governance Specialist, and founder of AI.MIRROR. Her research covers ten domains: the global AI governance crisis (Day 1), structural AI bias against women (Day 2), ML-based child criminality prediction (Day 3), the collapse of voluntary compliance in a military context (Day 4), LinkedIn’s algorithmic bias (Day 5), AI hallucination on non-Latin scr
GIGO and the Human Processor: A Chief Humanity Officer Review
This document is a peer review by Go Kian Tik — GKT (Mbah Hogi Bejo™) of Meriel B.’s 10-day research series. Meriel B. is an AI Ethics Researcher, AI Governance Specialist, and founder of AI.MIRROR. Her research covers ten domains: the global AI governance crisis (Day 1), structural AI bias against women (Day 2), ML-based child criminality prediction (Day 3), the collapse of voluntary compliance in a military context (Day 4), LinkedIn’s algorithmic bias (Day 5), AI hallucination on non-Latin scr
LLM Use, Cheating, and Academic Integrity in Software Engineering Education
Background: Cheating in university education is commonly described as context dependent and influenced by assessment design, institutional norms, and student interpretation. In software engineering education, programming oriented coursework has historically involved ambiguity around collaboration, reuse, and external assistance. Recently, large language models (LLMs) have introduced additional mediation in the production of code and related artifacts. Aims: This study investigates how software e
Impact of AI misinformation on diagnostic accuracy and confidence calibration in novice medical students
For novice medical learners, do the benefits of correct AI explanations outweigh the risks of plausible misinformation? In a randomized trial with 111 students, we found they do not. Our results reveal a significant and problematic asymmetry: misleading AI explanations significantly degraded diagnostic accuracy, while correct explanations offered no significant improvement over a no-explanation control. Misleading explanations reduced diagnostic accuracy and showed no evidence of confidence cali
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
Modeling plausible student misconceptions is critical for AI in education. In this work, we examine how large language models (LLMs) reason about misconceptions when generating multiple-choice distractors, a task that requires modeling incorrect yet plausible answers by coordinating solution knowledge, simulating student misconceptions, and evaluating plausibility. We introduce a taxonomy for analyzing the strategies used by state-of-the-art LLMs, examining their reasoning procedures and compari
AI to Learn 2.0: A Deliverable-Oriented Governance Framework and Maturity Rubric for Opaque AI in Learning-Intensive Domains
Generative AI is entering research, education, and professional work faster than current governance frameworks can specify how AI-assisted outputs should be judged in learning-intensive settings. The central problem is proxy failure: a polished artifact can be useful while no longer serving as credible evidence of the human understanding, judgment, or transfer ability that the work is supposed to cultivate or certify. This paper proposes AI to Learn 2.0, a deliverable-oriented governance framewo
Universe Routing: Why Self-Evolving Agents Need Epistemic Control
A critical failure mode of current lifelong agents is not lack of knowledge, but the inability to decide how to reason. When an agent encounters "Is this coin fair?" it must recognize whether to invoke frequentist hypothesis testing or Bayesian posterior inference - frameworks that are epistemologically incompatible. Mixing them produces not minor errors, but structural failures that propagate across decision chains. We formalize this as the universe routing problem: classifying questions into m
Retrieval-Feedback-Driven Distillation and Preference Alignment for Efficient LLM-based Query Expansion
Large language models have recently enabled a generative paradigm for query expansion, but their high inference cost makes direct deployment difficult in practical retrieval systems. To address this issue, a retrieval-feedback-driven distillation and preference-alignment framework is proposed to transfer retrieval-friendly expansion behavior from a strong teacher model to a compact student model. Rather than relying on few-shot exemplars at inference time, the framework first leverages two compl
Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments
Providing timely and individualised feedback on handwritten student work is highly beneficial for learning but difficult to achieve at scale. This challenge has become more pressing as generative AI undermines the reliability of take-home assessments, shifting emphasis toward supervised, in-class evaluation. We present a scalable, end-to-end workflow for LLM-assisted grading of short, pen-and-paper assessments. The workflow spans (1) constructing solution keys, (2) developing detailed rubric-sty
Learning from Child-Directed Speech in Two-Language Scenarios: A French-English Case Study
Research on developmentally plausible language models has largely focused on English, leaving open questions about multilingual settings. We present a systematic study of compact language models by extending BabyBERTa to English-French scenarios under strictly size-matched data conditions, covering monolingual, bilingual, and cross-lingual settings. Our design contrasts two types of training corpora: (i) child-directed speech (about 2.5M tokens), following BabyBERTa and related work, and (ii) mu
MRGEN: A Conceptual Framework for LLM-Powered Mixed Reality Authoring Tools for Education
Mixed Reality (MR) offers immersive and multimodal opportunities for education but remains difficult for teachers to author without technical expertise. We propose MRGEN, a conceptual framework for LLM-powered authoring tools to support teachers in creating MR learning activities that work on mobile devices (tablets and smartphones). MRGEN articulates three axes: Learning Objectives, MR Modality, and GAI Assistance. To validate our framework, we implemented a prototype based on the open-source M
Self-Regulated Personal Contracts as a Harm Reduction Approach to Generative AI in Undergraduate Programming Education
Students learning programming exercise agency in deciding when and how to use GenAI tools like ChatGPT. However, this agency is often implicit and shaped by deadline pressure and peer behavior rather than explicit and conscious learning goals. We designed a GenAI Contract grounded in harm reduction and self-regulated learning theory to scaffold intentional decision-making: students articulated personal learning goals, created usage guidelines, and reflected on alignment at strategic points acros
The Unlearning Mirage: A Dynamic Framework for Evaluating LLM Unlearning
Unlearning in Large Language Models (LLMs) aims to enhance safety, mitigate biases, and comply with legal mandates, such as the right to be forgotten. However, existing unlearning methods are brittle: minor query modifications, such as multi-hop reasoning and entity aliasing, can recover supposedly forgotten information. As a result, current evaluation metrics often create an illusion of effectiveness, failing to detect these vulnerabilities due to reliance on static, unstructured benchmarks. We
AI Detectors Fail Diverse Student Populations: A Mathematical Framing of Structural Detection Limits
Student experiences and empirical studies report that "black box" AI text detectors produce high false positive rates with disproportionate errors against certain student populations, yet typically theoretical analyses model detection as a test between two known distributions for human and AI prose. This framing omits the structural feature of university assessment whereby an assessor generally does not know the individual student's writing distribution, making the null hypothesis composite. Sta
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitation of rejection sampling. Standard methods treat the teacher as a static filter, discarding complex "corner-case" problems where the teacher fails to explore valid solutions independently, thereby creating an artificial "Teacher Ceiling" for the student. In this work, we propose Hindsight Entropy-Assisted Learning (HEAL), an RL-free framework designed to bridge this re
Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
Judging the novelty of research ideas is crucial for advancing science, enabling the identification of unexplored directions, and ensuring contributions meaningfully extend existing knowledge rather than reiterate minor variations. However, given the exponential growth of scientific literature, manually judging the novelty of research ideas through literature reviews is labor-intensive, subjective, and infeasible at scale. Therefore, recent efforts have proposed automated approaches for research
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
Medical image retrieval aims to identify clinically relevant lesion cases to support diagnostic decision making, education, and quality control. In practice, retrieval queries often combine a reference lesion image with textual descriptors such as dermoscopic features. We study composed vision-language retrieval for skin cancer, where each query consists of an image to text pair and the database contains biopsy-confirmed, multi-class disease cases. We propose a transformer based framework that l
A Text-Native Interface for Generative Video Authoring
Everyone can write their stories in freeform text format -- it's something we all learn in school. Yet storytelling via video requires one to learn specialized and complicated tools. In this paper, we introduce Doki, a text-native interface for generative video authoring, aligning video creation with the natural process of text writing. In Doki, writing text is the primary interaction: within a single document, users define assets, structure scenes, create shots, refine edits, and add audio. We
A Consensus-Driven Multi-LLM Pipeline for Missing-Person Investigations
The first 72 hours of a missing-person investigation are critical for successful recovery. Guardian is an end-to-end system designed to support missing-child investigation and early search planning. This paper presents the Guardian LLM Pipeline, a multi-model system in which LLMs are used for intelligent information extraction and processing related to missing-person search operations. The pipeline coordinates end-to-end execution across task-specialized LLM models and invokes a consensus LLM en
Interpretable Markov-Based Spatiotemporal Risk Surfaces for Missing-Child Search Planning with Reinforcement Learning and LLM-Based Quality Assurance
The first 72 hours of a missing-child investigation are critical for successful recovery. However, law enforcement agencies often face fragmented, unstructured data and a lack of dynamic, geospatial predictive tools. Our system, Guardian, provides an end-to-end decision-support system for missing-child investigation and early search planning. It converts heterogeneous, unstructured case documents into a schema-aligned spatiotemporal representation, enriches cases with geocoding and transportatio
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
We study the implicit bias of Sharpness-Aware Minimization (SAM) when training $L$-layer linear diagonal networks on linearly separable binary classification. For linear models ($L=1$), both $\ell_\infty$- and $\ell_2$-SAM recover the $\ell_2$ max-margin classifier, matching gradient descent (GD). However, for depth $L = 2$, the behavior changes drastically -- even on a single-example dataset. For $\ell_\infty$-SAM, the limit direction depends critically on initialization and can converge to $\m
AI Meets Mathematics Education: A Case Study on Supporting an Instructor in a Large Mathematics Class with Context-Aware AI
Large-enrollment university courses face persistent challenges in providing timely and scalable instructional support. While generative AI holds promise, its effective use depends on reliability and pedagogical alignment. We present a human-centered case study of AI-assisted support in a Calculus I course, implemented in close collaboration with the course instructor. We developed a system to answer students' questions on a discussion forum, fine-tuning a lightweight language model on 2,588 hist
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Continual Pre-Training (CPT). While effective for processing long sequences, this paradigm often disrupts original model capabilities, leading to performance degradation on standard short-text benchmarks. We propose LinearARD, a self-distillation method that restores Rotary Position Embeddings (RoPE)-scaled students through attention-structure consistency wit
IAML: Illumination-Aware Mirror Loss for Progressive Learning in Low-Light Image Enhancement Auto-encoders
This letter presents a novel training approach and loss function for learning low-light image enhancement auto-encoders. Our approach revolves around the use of a teacher-student auto-encoder setup coupled to a progressive learning approach where multi-scale information from clean image decoder feature maps is distilled into each layer of the student decoder in a mirrored fashion using a newly-proposed loss function termed Illumination-Aware Mirror Loss (IAML). IAML helps aligning the feature ma
AI Misuse in Education Is a Measurement Problem: Toward a Learning Visibility Framework
The rapid integration of conversational AI systems into educational settings has intensified ethical concerns about academic integrity, fairness, and students' cognitive development. Institutional responses have largely centered on AI detection tools and restrictive policies, yet such approaches have proven unreliable and ethically contentious. This paper reframes AI misuse in education not primarily as a detection problem, but as a measurement problem rooted in the loss of visibility into the l
Optimizing LLM Annotation of Classroom Discourse through Multi-Agent Orchestration
Large language models (LLMs) are increasingly positioned as scalable tools for annotating educational data, including classroom discourse, interaction logs, and qualitative learning artifacts. Their ability to rapidly summarize instructional interactions and assign rubric-aligned labels has fueled optimism about reducing the cost and time associated with expert human annotation. However, growing evidence suggests that single-pass LLM outputs remain unreliable for high-stakes educational construc
Eco-Bee: A Personalised Multi-Modal Agent for Advancing Student Climate Awareness and Sustainable Behaviour in Campus Ecosystems
Universities are microcosms of urban ecosystems, with concentrated consumption patterns in food, transport, energy, and product usage. These environments not only contribute substantially to sustainability pressures but also provide a unique opportunity to advance sustainability education and behavioural change at scale. As in most sectors, digital sustainability initiatives within universities remain narrowly focused on carbon calculations, typically providing static feedback that limits opport
The DSA's Blind Spot: Algorithmic Audit of Advertising and Minor Profiling on TikTok
Adolescents spend an increasing amount of their time in digital environments where their still-developing cognitive capacities leave them unable to recognize or resist commercial persuasion. Article 28(2) of the DSA responds to this vulnerability by prohibiting profiling-based advertising to minors. However, the regulation's narrow definition of "advertisement" excludes current advertising practices including influencer paid partnerships and brand promotional content that serve functionally equi
DARK: Diagonal-Anchored Repulsive Knowledge Distillation for Vision-Language Models under Extreme Compression
Compressing vision-language models for on-device deployment is increasingly important in clinical settings, but knowledge distillation (KD) degrades sharply when the teacher-student capacity gap spans an order of magnitude or more. We argue that, under such gaps, strict imitation of the teacher is a poor objective: much of the teacher's pairwise similarity structure reflects its own architectural biases rather than information a compact student can efficiently represent. We propose \textbf{Diago
Deterministic Preprocessing and Interpretable Fuzzy Banding for Cost-per-Student Reporting from Extracted Records
Administrative extracts are often exchanged as spreadsheets and may be read as reports in their own right during budgeting, workload review, and governance discussions. When an exported workbook becomes the reference snapshot for such decisions, the transformation can be checked by recomputation against a clearly identified input. A deterministic, rule-governed, file-based workflow is implemented in cad_processor.py. The script ingests a Casual Academic Database (CAD) export workbook and aggrega
AI agent in healthcare: applications, evaluations, and future directions
With the rapid advancement of large language model (LLM) technologies, AI agents have rapidly emerged in healthcare. This review traces the historical evolution and core characteristics of AI agents, and systematically examines their applications in assisted diagnosis, clinical decision support, medical report generation, patient-facing chatbots, healthcare system management, and medical education. We further analyze existing evaluation frameworks for AI agents in healthcare, focusing on key dim
Digital inequality in context: A socio-technical analysis of Arab students’ remote learning in Israel
The COVID-19 pandemic's shift to Emergency Remote Teaching (ERT) via platforms like Zoom created a global, real-world test for digitally mediated learning. This study provides an in-depth exploration of the multifaceted psycho-academic impacts on a particularly vulnerable population: Arab students in Israel, a minority group facing pre-existing socioeconomic and digital disparities. Through in-depth qualitative interviews with 30 students, I analyzed their lived experiences. To make sense of the
Evaluating AI-powered learning assistants in engineering higher education with implications for student engagement, ethics, and policy
As generative AI becomes increasingly integrated into higher education, understanding how students engage with these technologies is essential for responsible adoption. This study evaluates the Educational AI Hub, an AI-powered learning framework, implemented in undergraduate civil and environmental engineering courses at a large R1 public university. Using a mixed-methods design combining pre- and post-surveys, system usage logs, and qualitative analysis of students' AI interactions, the resear
Artificial intelligence in higher education: a systematic review of its impact on student engagement and the mediating role of teaching methods
Introduction Artificial Intelligence (AI) is increasingly integrated into higher education to personalize instruction and support student engagement. However, the mediating role of teaching methods in this process remains underexplored. Methods This systematic review analyzed 73 peer-reviewed articles published between 2015 and early 2025, retrieved from Scopus and Web of Science, following PRISMA guidelines. Studies were screened based on predefined inclusion criteria and coded using a structur
Evaluating the accuracy and reliability of AI content detectors in academic contexts
The rapid adoption of generative AI (GenAI) in higher education has intensified concerns about academic integrity, particularly for institutions serving English as a Foreign Language (EFL) learners. AI content detectors such as Turnitin and Originality are now widely used to identify potential misuse of GenAI in student writing, yet their accuracy, consistency, and fairness remain to be proven. This study evaluates the reliability of these two commercial detectors using a balanced dataset of 192
Development and evaluation of artificial intelligence literacy training for teacher education students
Abstract Teacher education students play double role as present learners and future educators. Hence, they need targeted training to navigate the growing influence of Generative Artificial Intelligence (GenAI) on teaching, learning and professional identity. However, existing artificial intelligence (AI) literacy programmes predominantly emphasize technical AI knowledge and pre‐GenAI tools or are offered by GenAI platforms that focus on their own technologies' features, thereby lacking pedagogic
Exploring the role of agentic AI in fostering self-efficacy, autonomy support, and self-learning motivation in higher education
Introduction: Rapid adoption of Artificial Intelligence (AI) in learning has revolutionized learners' engagement but comprehension of psychological and technological drivers of successful AI-enabled learning remains scarce. This research investigates how students' perceived agency of AI, usefulness, ease of use, trust, autonomy supporting, and self-efficacy collectively impact students' self-learning behavior and motivation. Based on Technology Acceptance Model (TAM), Social Cognitive Theory (SC
Teachers’ artificial intelligence (AI) literacy: an exploratory study
Abstract This study explores variables associated with teachers’ Artificial Intelligence (AI) literacy, a key competency for effective and responsible AI integration in education. A total of 270 teachers completed an online survey including measures of AI literacy, AI acceptance, computational thinking, AI anxiety, and digital divide. Results revealed that all AI acceptance variables were positively associated with AI literacy, with hedonic motivation and willingness to use AI emerging as the st
Digital Ecosystems, Children, and Adolescents: Policy Statement
Digital media, including television, the internet, social media, video games, and interactive assistants, form the digital ecosystem. When this digital ecosystem is designed with children's unique developmental needs in mind, it can support learning and well-being. In contrast, digital ecosystems that prioritize engagement and commercialization often encourage prolonged use, which in turn can displace healthy behaviors (eg, movement behaviors, sleep), and contribute to negative outcomes. This po
Development and validation of the AI dependence scale for Chinese undergraduates and a preliminary exploration
Introduction: With the proliferation of generative artificial intelligence (AI) in higher education, student overreliance has become a growing concern, potentially undermining critical thinking and autonomous learning. To address the lack of a comprehensive measurement tool, this study developed and validated the AI Dependence Scale (AIDep-22), a new instrument designed to assess this phenomenon across four hypothesized dimensions: emotional dependence, functional dependence, cognitive dependenc