Research · RSS feed
New papers on fairness, safety, alignment and governance.
A comparative and integrative mapping of thematic trajectories, the digital turn, and future research frontiers in telecommunications policy
Publication date: September 2026 Source: Telecommunications Policy, Volume 50, Issue 8 Author(s): Menglan Luo, Yong Jiang, Yi-Shuai Ren
How demography shapes the macroeconomic returns to automation
Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Izabella Kuncz, Petra Németh, Eszter Szabó-Bakos
ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts
Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained by three fundamental bottlenecks. These encompass the massive memory bandwidth required to fetch non-contiguous expert weights, the non-deterministic scatter-gather traffic generated by input-dependent token routing, and the tail-latency dependency imposed by synchronous expert output aggregation. To address these chal
“You Might Like Each Other”: Mass Customization and the Structuring of Sociality Among Chinese Generation Z on the Soul App
Social Media + Society, Volume 12, Issue 3, July-September 2026. We employed the framework of mass customization to investigate how the Soul app, a leading interest-based social platform, is reshaping sociality among Chinese Generation Z. Six months of participatory observation, from December 2024 to June 2025, in Soul’...
Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning
LLM constraint reasoners are often evaluated near the random-SAT phase transition, confounding density and solver hardness. We test instance-level transfer while near-matching clause density. At aligned size bins, with near-matched density and matched maximum clause width, we compare proof-hard expander-Tseitin and proof-easy ladder-Tseitin formulas, pigeonhole anchors, and density-mismatched controls. Theory separates their resolution hardness; a solver-specific Glucose mean-conflict proxy diff
Federated Lightweight Intrusion Detection in Drone Swarms with Knowledge Distillation
Drone swarms are increasingly deployed in critical applications such as surveillance, disaster response, and infrastructure monitoring. However, their reliance on open communication channels and their limited computational resources make them vulnerable to a wide range of cyber-threats. There is a growing interest in intrusion detection systems (IDS) specifically designed for drone environments and operations. However, the conventional solutions including Machine Learning (ML)-based approaches r
The value of contact in legged locomotion: a survey of sensing channels, artificial intelligence and control
Legged robots traverse unstructured terrain through brief, intermittent foot–ground contacts whose support conditions are difficult to perceive and predict in real time. In such regimes, haptic feedback provides early and trustworthy evidence of traction limits, partial support, and incipient slip. This structured survey asks two questions: first, what locomotion-relevant contact evidence can be acquired and preserved under real deployment constraints; and second, how that evidence is translated
The shared blind spot: why diverse AI governance approaches fail for the same reason
AI governance instruments are proliferating, and so are their difficulties. Across major jurisdictions and international bodies, reform efforts built on substantially different premises encounter a recognizably similar pattern of failure. I argue that anticipatory regulatory governance rests on three operational premises —categorical stability, epistemic accessibility, and manageable pace—and that AI’s emergence, opacity, and velocity violate all three in compound. These premises form a distinct
Counterfactual Shapley Credit Assignment
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reward reweighting, frequently fail to attribute properly between an agent's policy (skill) and environmental stochasticity (luck). A principled approach to CAP must isolate the true causal drivers of observed outcomes from spurious correlations and environmental randomness. We introduce
Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each scholar's record by hand takes staff an estimated 15 hours and does not scale to a full cohort. An artificial intelligence (AI) agent could serve as a tool to gather scholar data across platforms and disciplines. Methods. We built a human-in-the-loop AI agent that assembles a dossier of sourced evidence for each scholar and drafts one-sentence Translational Sc
Migration and Displacement in Crisis: Protection, Policy, and Practice
This course examines refugees, migrants, stateless people, and internally displaced persons in contemporary crisis settings. It explores how conflict, political violence, disasters, climate stress, ...
Distilled Reinforcement Learning for LLM Post-training
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture how characters and worlds evolve together over time. To address this, EvolvingWorld models literary simulation as a long-horizon process where characters interact, scenes progress, and character and world states are persistently
A Red Line and Oversight Framework for Government AI Contracts
A Method for Learning Value Systems in Generative AI
Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems in order to align their decisions with ours. As such representations are difficult to elicit, value learning seeks to infer them by observing human behaviour. This work addresses the lack of grounded value learning methods in generative AI: existing approaches typically replicate human preferences without awareness of the multidimensional structure of value
Trace-Based On-Policy Distillation for Masked Diffusion Language Models
Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervised fine-tuning (SFT) for dLLMs requires dense but often off-policy masked states, while reinforcement learning (RL) relies on sparse rewards or value modeling. This paper proposes \textbf{trace-based on-policy distillation (TOPD)}, a teacher-supervised framework that transfers reasoning ability to a target dLLM without
Endogenous Alignment
MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning
Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities. However, training multimodal models faces two main obstacles. First, collecting large-scale, well-aligned paired multimodal datasets is often impractical, making end-to-end multimodal training difficult. Second, existing multimodal representations frequently entangle information shared across modalities with modality-specific information, hindering interpretability and contro
Generative AI and the uneven disruption of freelance labor: Exposure, diffusion, and adaptation under constraint
Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Zeljko Tekic, Yury Omelchenko, Arina Dunaeva
Making privacy usable: Bridging privacy research and practice through guidelines
Publication date: Available online 16 July 2026 Source: International Journal of Human-Computer Studies Author(s): Shirlei Aparecida de Chaves, Fabiane Barreto Vavassori Benitti
AI Literacy-Related Domains and AI-TPACK Readiness Among Preservice Mathematics Teachers: A Factor-Informed Structural Equation Modelling Study
Publication date: Available online 16 July 2026 Source: Computers and Education: Artificial Intelligence Author(s): Moeketsi Mosia, Fadip Audu Nannim, Felix Egara
Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis
Publication date: Available online 16 July 2026 Source: Computers and Education: Artificial Intelligence Author(s): Kamila Misiejuk, Sonsoles López-Pernas, Eduardo Araujo Oliveira, Brendan Eagan, Mohammed Saqr
Privacy Cost as Equity Input: A Group Fairness Criterion for Differentially Private Machine Learning
Differential privacy (DP) is increasingly deployed to limit membership inference risk in machine-learning systems. Prior work has shown that DP-SGD can widen accuracy disparities across demographic groups, but this framing treats fairness as a purely outcome-side concern. We argue that privacy cost, the information leakage borne by each group, is itself a form of harm, and adopt a compensatory-fairness framework in which a group that involuntarily bears greater privacy exposure is owed proportio
Should we benchmark conceptual capabilities using judgment prediction tasks?
Hindsight: Similarity-Based Analytics for Mars Rover Drive Retrieval
While Mars rover operators plan drives across hazardous Martian terrain and diagnose unexpected faults, the necessary information is distributed across separate systems and often reconstructed through manual correlation and memory. To address this challenge, we partnered with Mars rover operators at the NASA Jet Propulsion Laboratory to introduce Hindsight, a visual analytics system that unifies previously disparate rover drive data into a single workspace for search, comparison, and investigati
How Formerly Incarcerated People Envision Technologies for Prison Parole
AI-driven algorithms and automated tools are increasingly embedded in the correctional landscape, shaping parole eligibility,release decisions, and surveillance. These tools are also often framed as objective, inevitable solutions to inefficiency andbias. Yet, these computational systems are rarely designed with input from justice-impacted individuals, which means theymight fail to address the real needs of incarcerated people. To address this gap, we surveyed 31 formerly incarcerated peopleabou
Implementing Supported Digital Enhanced Cognitive Behavior Therapy for Binge Eating Disorder in Routine Care: Mixed Methods Service Evaluation
Background: Binge eating disorder (BED) is highly prevalent and impairing; yet, the UK national guideline–recommended first-line treatment of guided self-help (ie, supported program-led interventions in which content is delivered by the program with brief support) remains underused in routine National Health Service (NHS) care. Digital delivery offers a scalable approach, but evidence from real-world NHS settings is limited. Objective: This real-world evaluation aimed to pilot a supported digita
From Spatial AI to Clinical Readiness: Lessons From Augmented World Expo USA 2026
Group Entropy-Controlled Policy Optimization
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on mixtures of heterogeneous tasks, which induce distinct entropy regimes under the same policy, making global or token-level entropy regulation insufficient to corresponding heterogeneous needs of exploration. This heterogeneity further makes GRPO-style normalized advantages i
Environment-free Synthetic Data Generation for API-Calling Agents
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method genera
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the NL2Pipeline gap. To bridge it, we introduce DataFlow-Harness, a platform that guides an LLM agent to construct platform-native directed acyclic graphs (DAGs) through typed, incremental mutations rather than free-form scripts. The platform combines DataFl
Can Multimodal Large Language Models Understand OCT?
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we
CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii) state-agnostic divergence scheduling, where time-only forward/reverse-KL interpolation ignores the student's coverage state; and (iii) binary reward sparsity, where pass/fail signals discard information from partially correct traces. We p
How Pandemics Have Reshaped the Respiratory Virus Data Landscape in Europe: Scoping Review
Background: Acute respiratory infections caused by influenza, respiratory syncytial virus (RSV), and SARS-CoV-2 remain a major public health challenge in Europe. Although surveillance systems for these pathogens are well established, the past 2 decades have seen a rapid diversification of data streams supporting surveillance and research. This expanding data landscape, combined with fragmentation across institutions, sectors, and countries, may limit timely evidence synthesis and effective publi
Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum
For a second time, the android robot Andrea was set up at a public museum in Germany for six consecutive days to have conversations with visitors, fully autonomously. Building on previously gathered qualitative results, the robot was now capable of engaging in multi-lingual conversation with the visitors about the museum context. The robot was prepared with context information about the museum in general and its surrounding exhibits this time. The robot featured a slightly artificial sounding vo
Announcing the Corrigibility Research Fund
Cluster-Aware Matching via Laplacian Optimal Transport
In many applications of matching, the point clouds to be matched are not merely unstructured sets of points but rather samples from distributions with an intrinsic cluster structure. In such cases, as individual points are often interchangeable within a coherent region, finding a robust region-to-region alignment is more desirable than establishing a precise point-to-point correspondence. To this end, we propose a novel approach for cluster-aware matching based on Laplacian Optimal Transport (La
When Does Muon Help Agentic Reinforcement Learning?
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The
When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an information bottleneck perspective on elucidating the differences between MAS and SAS. Specifically, our key observation is that a SAS accumulates its full reasoning trace in one shared context, while a MAS uses isolated local contexts connected by bounde