Research · RSS feed
New papers on fairness, safety, alignment and governance.
Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments
arXiv:2607.18874v1 Announce Type: cross Abstract: Using Unmanned Aerial Vehicle (UAV) for urban sensing has emerged as a powerful paradigm to monitor the status of the city, e.g., air quality and noise levels, through agile aerial crowdsourcing. Despite this potential, existing UAV-based sensing approaches overlook environmental disturbances like wind that drastically impact drone velocity and energy efficiency. Consequently, directly applying existing methods to this joint delivery and sensing
From Operations to Elderly Care Outcomes: A Thematic Review of Industrial Engineering and Decision-Support Approaches
arXiv:2607.19075v1 Announce Type: cross Abstract: The rapid growth of the global aging population presents severe challenges to healthcare systems, necessitating efficient, equitable, and patient-centered care models. While Industrial Engineering and Operations Research (OR) provide robust optimization and decision-support tools to address these multidimensional complexities, current applications often remain fragmented. This paper presents a thematic review of 30 seminal studies at the intersec
Global Automation Atlas
arXiv:2605.17086v2 Announce Type: replace-cross Abstract: Automation can displace or complement labour, but this need not be constant across economies. Existing exposure measures typically assign fixed scores to tasks or occupations and capture cross-country variation through employment structure. Here we show that feasible automation depends jointly on task content and country-level conditions. We use a large language model to classify 18,797 work tasks in 124 economies by exposure, labour marg
The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation
arXiv:2605.22840v2 Announce Type: replace-cross Abstract: How much thinking can a civilisation do? Kardashev ranked civilisations by the energy they command. This paper borrows his ladder and asks how much machine cognition each rung could support. The arithmetic is deliberately simple. A civilisation has some total power. Only a fraction of that can be spared for computing, and each joule spent buys computation at whatever efficiency the hardware of the day has reached. The product of the three
Did Alice Do Wrong? Cross-Cultural Differences in Student Perceptions of Generative AI Use in University Computing Education
The rise of generative AI (GenAI) in higher education has prompted urgent debates surrounding academic integrity and ethical use. This study examines cross-cultural differences in student perceptions of GenAI use, comparing responses from students at Canadian and South Korean universities. Using a scenario-based survey administered in Fall 2024, we analyzed how students judged the ethicality and rule compliance of AI-assisted coding practices. Results reveal that Canadian students were consisten
A Drift Stable Quantum Federated Learning for Intelligent Services
Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aware intelligent services. Intelligent services in this context refer to privacy-sensitive distributed decision systems, such as fraud detection and genomic classification, where reliable and fair client-level learning is as important as the accuracy of the aggregate model. However, heterogeneous client data and noisy quantum optimization often caus
Associations Between Support-Seekers' Cross-Community Interactions and Their Engagement with Received Comments in Online Health Communities
Support-seeker' active engagement with received comments, e.g., showing positive sentiment and willingness to improve in the replies, can indicate the success of online health communities (OHCs). Their participation in other communities may correlate with their engagement in OHCs but remains under-explored. This paper analyzes 26, 725 seekers' behaviors in the other 40, 479 communities and their associations with seekers' engagement with received comments under their 78, 501 posts in 30 Baidu Ti
Decoupling what, how, and when for observing decision-making context in autonomous robots
This paper presents an approach that decouples what to observe, how to observe it, and when observations are required for decision-making in autonomous robots. Situation awareness is essential for efficient and reliable autonomous robot operation, but despite advances toward parallelizing perception and action, key challenges remain in making perception aware of the current context and ensuring observability during action execution. To address this, we explicitly model the decision-making contex
Large language models and multimodal AI for mental health: a systematic review of early diagnosis and monitoring
Mental health disorders (e.g., depression, anxiety, post-traumatic stress disorder (PTSD), bipolar disorder) represent a pressing global challenge, and early diagnosis with continuous monitoring is critical for effective intervention. However, traditional diagnostic methods, relying on patient self-reports and clinical interviews, are subjective and often miss subtle early warning signs, a problem compounded by stigma and limited access to care. In response, recent advances in artificial intelli
LLM as a law professor: having a large language model write a commentary on freedom of assembly
In many jurisdictions, academia is at the service of legal practice. Law professors write commentaries that summarize the state of the art of doctrine, chiefly of jurisprudence. In the spirit of a proof of concept, using the guarantee of freedom of assembly in the European Convention on Human Rights, we show that this task can be completely outsourced to large language models. Using standard NLP metrics and an LLM as a judge approach, we develop an evaluation pipeline that works without costly h
Climate-control ChatGPT for EFL writing and AI literacy
ChatGPT is increasingly used in EFL writing, yet applied linguistics lacks robust methods for measuring whether human-AI dialogue improves learners’ AI literacy, harm-sensitive reasoning, and independent argumentative writing. This mixed-methods quasi-experimental study examines whether climate-control ChatGPT, a prompted interactional design that requires learners to justify, qualify, counterexample, and re-author their own claims, produces stronger gains than standard ChatGPT use and no-ChatGP
Correction: Performance of large language models in neonatal resuscitation assessments versus healthcare providers: an exploratory study
On the fragility of neural architecture search: the role of overfitting and task complexity in medical image analysis
IntroductionNeural Architecture Search (NAS) effectively automates Deep Learning pipeline design but is prone to validation overfitting when applied to complex tasks, such as medical image analysis. To mitigate this and enhance generalization, researchers frequently integrate Deep Ensemble Learning (DEL) and data augmentation into the NAS workflow. However, the assumption that these methodologies do not negatively interfere in high-overfitting scenarios remains unproven.MethodsWe evaluated NAS,
How U.S. Federal Artificial Intelligence (AI) policy is shaping agrifood systems: an integrative review
Artificial intelligence (AI) is increasingly shaping how agrifood systems function in the United States, yet the role of federal policy in guiding its use, oversight, and broader consequences is still not well defined. This study explores how current U.S. federal AI policies support-or limit-the advancement of agrifood systems by synthesizing evidence from publicly available policy documents. Using a focused search strategy and qualitative content analysis, we reviewed nine federal policy docume
Application of artificial intelligence models in the identification of severe scrub typhus
This retrospective study enrolled 492 patients with scrub typhus in Jiangmen from 2013 to 2025. Clinical and laboratory data were analyzed using univariate logistic regression and LASSO regression to identify risk factors for severe illness. Seven machine-learning models, including logistic regression, support vector machine, random forest, XGBoost, Naive Bayes, k-nearest neighbor, and decision tree, were constructed and externally validated. Variable importance was ranked using SHAP analysis. M
Understanding Developer Pain Points in Federated Learning: Insights from Stack Overflow and GitHub
Federated Learning (FL) enables collaborative model training without centralizing raw data, but building and operating FL systems remains difficult due to distributed execution, rapidly evolving frameworks, and privacy and governance requirements. In this paper, we present an empirical study of FL developer challenges by independently analyzing 495 Stack Overflow posts and 9,116 GitHub issues and pull requests from 92 FL-related projects. Using BERTopic-based topic modeling and difficulty indica
Deep Shape Regression for Planar Curves with Multimodal Covariates
The shape of a planar curve is the geometric information that remains once translation, rotation, scale and reparametrisation are removed and is of interest in many health applications, e.g. in neuroimaging. We propose a deep shape regression model for open planar curves that admits multimodal and high-dimensional covariates. Representing curves as complex-valued functions, we show that the conditional full Procrustes mean is the leading eigenfunction of the conditional covariance. To estimate t
Communication Barriers in Patient-Provider Interactions in Health Care: Scoping Review
Background: Effective communication is crucial for high-quality health care, but systemic barriers still disrupt patient-provider interactions. Research shows that communication failures are a major reason for preventable medical errors. These issues are linked to around 30% of malpractice claims and over 1744 deaths each year in the United States. In addition, hospitals lose about US $12 billion each year because of miscommunication. Despite the critical nature of this issue, the literature rem
Examining User Behavior and Cognitive Biases in Personal Password Security
Despite increasing awareness of cybersecurity risks, users continue to engage in insecure password practices, such as reusing passwords, choosing weak credentials, and neglecting security recommendations. The study explores the behavioral and cognitive factors that influence password decision-making by integrating insights from behavioral economics, particularly hyperbolic discounting, status quo bias, and present bias. We conducted a survey to analyze how people create, store and manage their p
End-to-End Differential Privacy in Training Deep Neural Network Classifiers
Differentially private machine learning enables model training on sensitive data while ensuring that individual data is unlikely to be recoverable from the parameters of the resulting model. However, existing work often privatizes both training inputs and their labels, and these protections may be conservative when labels are public or can be safely made public. Therefore, in this work we propose a novel private training framework that instead privatizes training inputs while keeping labels publ
Effects of a Short Mobile Intervention on Digital Health Literacy in Adolescents and Teachers: Randomized Controlled Trial
Background: Digital health literacy is an essential skill for processing health-related information in today’s technology-driven society. Recent literature highlights deficits in digital health literacy among adolescents and legislative initiatives to secure its promotion have been introduced (eg, in the German Social Code Book V). However, few interventions target adolescents; existing programs often overlook critical aspects like graph literacy or lack applicability in schools and rigorous sci
Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning
Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible. We introduce Trace, a taxonomy-guided environment for multidomain visual reasoning. Trace factorizes task construction into a scene grammar and an executable task program, separating visual realization from answer computation. A sh
Self Gradient Forcing: Native Long Video Extrapolation
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, conflict, complementarity) and unable to answer what makes one document set better than another. To address these issues, we propose a complete evaluate-diagnose-optimize framework
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, whil
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framework supports structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture, and automatic bounding box annotation using engine-le
SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to predict the grasp directly with limited spatial awareness, or train the VLM together with the grasping model, which requires significantly more data and compute. These limitations impede performance and have prevented scaling to mult
LLMs Get Lost in Evolving User Intent
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user inten
Robostral Navigate
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor acr
ReferTrack: Referring Then Tracking for Embodied Visual Tracking
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections. To address this, we introduce ReferTrack, a referring-then-tracking para
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions the model can take, fields are its state, docstrings are its prompts, and its type annotations are contracts. A method whose code body consists of "..." is completed at runtime by an
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual en
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematica
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue: the cost lands on the researcher at every iteration. Molt is a PyTorch-native training framework built to keep that cost small: a codebase compact and clean enough for a researcher to hold in their head, and for an AI coding assistant to read and reas
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations
Multimodal Speaker Verification as a Threat to Speaker Anonymization
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across a
How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
In RLHF pipelines, reward scoring blocks policy updates. Slow scoring bottlenecks the entire loop, since no update runs until every rollout gets a score. And yet most setups just default to PyTorch eager mode or torch.compile, no one checks if that's actually fastest. Scoring itself is small. Rollout generation eats far more of a typical RLHF step. But scoring and generation fight over the same CPU and GPU resources, so a faster scoring engine doesn't shrink step time on its own. It mainly frees
Trustworthy Privacy-Preserving Multimodal Federated Learning for Personalised Breast Cancer Prediction
Federated learning has emerged as a potential solution to privacy concerns associated with using sensitive health data for training predictive models, particularly in personalised cancer care. This research investigates whether federated learning can support the development of robust models for predicting tumour progression in breast cancer patients while addressing four critical deployment pillars: transparency, scalability, security, and fairness. This study evaluates a federated learning fram
A Supervised Fine-Tuned Large Language Model for Lifestyle Management in Patients With Prostate Cancer: Development and Evaluation Study
Background: Lifestyle interventions for patients with prostate cancer have been shown to improve treatment adherence and quality of life. However, there remains a lack of large language models (LLMs) capable of delivering individualized and professional lifestyle recommendations under clearly defined medical safety boundaries and controlled evidence sources. Objective: This study aimed to develop and evaluate a supervised fine-tuned LLM—PCaPLMM_SFT (Prostate Cancer Patient Lifestyle Management M
AI-Powered Simulation for Nursing Education: Mixed Methods Systematic Review
Background: Traditional simulation-based nursing education is often constrained by high costs, resource intensity, and limited scalability. AI-powered simulations offer dynamic, scalable, and personalized alternatives. However, the empirical evidence regarding their pedagogical effectiveness and learner acceptance remains fragmented. Objective: This study aimed to systematically evaluate and synthesize evidence on the effectiveness and learner perceptions of AI-powered simulations in nursing edu