13:05 UTC

LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods. Here, foundation models are pre-trained on mixtures of complex clinical data modalities, useful for various downstream tasks. Existing works often utilise Electronic Health Records (EHR) to provide rich and diverse patient observations to train clinical foundation models. Howe
arXiv cs.LG 15d ago Healthcare

Describing a National Chatbot Deployed by the Ministry of Health in Malawi During the COVID-19 Pandemic: Retrospective Data Analysis

Background: Malawi was a pioneer among African countries in implementing a coordinated, government-led effort to streamline COVID-19 support using digital health tools. In response to the pandemic, a COVID-19 WhatsApp chatbot was developed to support the public with information, symptom reporting, and service navigation during the pandemic. Objective: This study describes the national deployment, functionality, and use patterns of the WhatsApp chatbot during the COVID-19 pandemic in Malawi. Meth
JMIR (Journal of Medical Internet Research) 15d ago Healthcare

DSWorld: A Data Science World Model for Efficient Autonomous Agents

Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects of data science operations before real execution. In this paper, we introduce the concept of Data Science World Model, which model the data science execution environment by predicting environment state transitions conditioned on current workflow sta
HuggingFace Daily Papers 15d ago Agents & autonomyEnvironment

Recursive Harness Self-Improvement

Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a ta
HuggingFace Daily Papers 15d ago Jobs & economyAgents & autonomy

RecGPT-V3 Technical Report

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior mo
HuggingFace Daily Papers 15d ago Agents & autonomy

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloa
HuggingFace Daily Papers 15d ago Agents & autonomy

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 in
HuggingFace Daily Papers 15d ago Agents & autonomy

When Does Muon Help Agentic Reinforcement Learning?

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The
HuggingFace Daily Papers 15d ago RegulationAgents & autonomy

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneous action can lead to irreversible consequences. Existing safety mechanisms are primarily reactive, lacking the ability to assess risks before execution. In this paper, we introduce SeerGuard, a consequence-aware safety framework designed to mitigate these risks through pre-execution instruction-level screening and acti
HuggingFace Daily Papers 15d ago Agents & autonomy

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos cove
HuggingFace Daily Papers 15d ago Regulation

Interactive Training 2: Auditable Control Plane for Live Model Training

Experiment trackers show how training is progressing, but changing a live run still usually requires trainer-specific code. We present Interactive Training 2, an open-source control plane for steering training through a shared protocol. Training applications declare which settings and actions they expose, humans and automated controllers submit requests through the same interface, and the training loop validates and applies them at safe control points. A customized Aim workspace combines live me
HuggingFace Daily Papers 15d ago Transparency

Translation and Psychometric Validation of the Amharic eHealth Literacy Questionnaire: Cross-Sectional Study

Background: eHealth interventions have demonstrated potential to address challenges related to health and the health care system in low- and middle-income countries. To effectively leverage eHealth in supporting health care in Ethiopia, the assessment and development of the eHealth literacy of patients are essential. Objective: This study aimed to translate the eHealth Literacy Questionnaire (eHLQ) to Amharic and assess its psychometric properties. Methods: A systematic process of translation, i
JMIR (Journal of Medical Internet Research) 15d ago Healthcare

Frontiers of AI in Negotiation and Conflict Management

Hannah Riley Bowles Appointment Co-Director, Women and Public Policy Program Roy E. Larsen Senior Lecturer in Public Policy and Management Mailing Address 1 Harvard Kennedy School Mailing Address 2 79 ...
Harvard Kennedy School 15d ago RegulationChildren & education

Assessing Learning Processes with Multimodal Data in Virtual Reality Learning Environments

Assessing learning in virtual reality (VR) environments typically relied on traditional pre-post content retention tests, revealing little about the process of learnng within such immersive environments. Multimodal data from player activity in VR is promising to better measure learning processes and higher-order skills, but little research in the learning sciences has explored how such data can be combined to provide meaningful measures. To address this, we explored multimodal sources of data fr
arXiv cs.HC 15d ago Environment

Clinical Audit Logs as Multi-Axial Traces of Care Delivery

Electronic health record audit logs record timestamped actions through which clinical work is carried out. Generated as operational metadata, they now support research on clinician effort, patient outcomes, care-team coordination, and workflow structure. This Perspective explains that breadth by articulating audit logs as multi-axial event streams and drawing implications for representation learning, evaluation, and governance. Each logged action belongs simultaneously to multiple clinically mea
arXiv 15d ago RegulationHealthcare

Development and Clinical Evaluation of a Large Language Model–Based System for Generating Patient-Friendly Echocardiography Reports: Two-Stage Retrospective Validation and Prospective Survey Study

Background: Standard echocardiography reports use complex terminology, limiting patient comprehension and exacerbating preconsultation anxiety. Large language models (LLMs) can transform technical data into patient-friendly narratives by incorporating longitudinal comparisons with prior examinations. Objective: This study aims to develop an LLM-based patient-friendly echocardiography reporting system and evaluate its professional safety, patient comprehension, and impact on short-term anxiety. M
JMIR (Journal of Medical Internet Research) 15d ago Healthcare

Application of Just-in-Time Adaptive Interventions in Dietary Health Management: Systematic Review

Background: Just-in-time adaptive interventions (JITAIs) use real-time data to deliver personalized support at moments of heightened need and may improve dietary behaviors in real-world settings. Objective: The aim of this study is to systematically review the application, characteristics, and effectiveness of JITAIs in dietary health management. Methods: We included human studies evaluating JITAIs-based dietary interventions delivered through digital platforms that used real-time or near–real-t
JMIR (Journal of Medical Internet Research) 15d ago Healthcare

An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

Large language models (LLMs) are increasingly embedded in clinical and population health workflows, including conversational agents such as health chatbots. As chatbots evolve from rule-based approaches to hybrid and LLM-enabled designs, risks and concerns about deployment readiness shift. Unlike rule-based chatbots, LLM outputs can be unpredictable, error-prone, and difficult to validate with traditional evaluation methods. Public health teams integrating customized LLMs into interventions face
JMIR (Journal of Medical Internet Research) 15d ago HealthcareAgents & autonomy

Consensus Statement on Digital Health and Attention-Deficit/Hyperactivity Disorder by the European Network for ADHD (EUNETHYDIS): Modified Delphi Study

Background: Digital technologies are becoming an important part of health care, including for individuals with attention-deficit/hyperactivity disorder (ADHD). Digital health innovations present valuable opportunities to provide flexible and tailored support for their diverse needs, along with significant challenges. Attentional, organizational, and motivational characteristics associated with ADHD may affect how individuals engage with digital tools. Potential risks include additional access ba
JMIR (Journal of Medical Internet Research) 15d ago Healthcare

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturb
arXiv cs.AI 15d ago RegulationAgents & autonomy

SceneBind: Binding What and Where Across Vision, Audio and Language

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explic
arXiv 15d ago

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 in
arXiv cs.AI 15d ago Agents & autonomy

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can become trapped in repetitive loops, wasting search budgets and ultimately compromising the quality and completeness of the final output. We introduce SearchOS, a system-level multi-age
arXiv cs.AI 15d ago Agents & autonomy

AutoSynthesis: An agentic system for automated meta-analysis

Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quantitative evidence synthesis remains largely manual and difficult to scale. Here, we introduce AutoSynthesis, an end-to-end multi-agent system for automated meta-analysis. Given a research question in natural language, AutoSynthesis formulates a search strategy, retrieves scientific literature, screens candidate studies, assesses full-text eligibility, extracts
arXiv 15d ago RegulationChildren & education

Publicly Accessible Large Language Model Responses to Frequently Asked Questions About Spondylodiscitis: Preliminary Expert Evaluation

Background: Patients increasingly use large language models (LLMs) to obtain medical information, but the quality of LLM-generated information on complex spinal infections such as spondylodiscitis remains uncertain. Existing evaluations in spine surgery have mainly addressed degenerative conditions or surgical procedures, and disease-specific data for spondylodiscitis are limited. Objective: This preliminary study evaluated spine surgeons’ ratings of single-turn LLM responses to 10 author-curate
JMIR (Journal of Medical Internet Research) 15d ago Healthcare

In-Place Tokenizer Expansion for Pre-trained LLMs

A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into many more tokens per word, which can raise latency, compute, and energy consumption for users of those languages. Cloud models can afford a broad vocabulary because the embedding and LM-head matrices are a small fraction of their parameters. On a compact model those matric
arXiv cs.AI 15d ago Environment

Divergent Gaze Patterns in Artistic Viewing: Spatial and Temporal Signatures of Attention Across Autistic Individuals, Artists, and Neurotypical Observers

How different populations visually explore artworks bears on cognitive science and on accessibility design, yet most eye-tracking work in autism has used social scenes rather than art, and has analysed where the eyes land while ignoring when and in what order. We present a comparative free-viewing study across three groups, autistic adults (ASD), trained artists, and neurotypical observers, who each viewed 30 paintings for 15s. We introduce a directed, metric-grounded framework that compares gro
arXiv cs.HC 15d ago Privacy

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-state direction analysis and random-split null tests, we show that content danger (CD) and physical danger (PD) form separable signals in LLM representations across Qwen2.5-3B/7B/14B
arXiv cs.AI 15d ago Agents & autonomy

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer t
arXiv cs.AI 15d ago Safety & alignment

BadWAM: When World-Action Models Dream Right but Act Wrong

World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluatin
arXiv cs.LG 15d ago Safety & alignmentAgents & autonomy

Can giant space mirrors boost green energy on Earth? A start-up aims to find out

A risky plan to turn night into day is one step closer to reality. Last week, US officials approved a mission to launch a giant mirror into space, where the device will reflect sunlight onto shadowed ...
Nature Machine Intelligence 15d ago Environment

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health research. In depression-related datasets, labels are often assigned without structured evidence, symptom-level justification, or traceable alignment with the criteria of the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition, Text Revision (DSM-5-TR), limiting both transparency and downstream model interpretability. We propose a self-evolving, ex
arXiv cs.HC 15d ago Safety & alignmentHealthcare

Mask-Aware Policy Gradients for Diffusion Language Models

Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling only the token predictions, ignoring the order in which positions are unmasked during generation. We observe that MDLM generation involves two decisions at each step: what tokens to place at each maske
arXiv cs.AI 15d ago Regulation

Plover: Steering GUI Agents through Plan-Centric Interaction

Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users' ability to inspect, supervise, or correct system behavior. We present Plover, a pla
arXiv cs.AI 15d ago Jobs & economyAgents & autonomy

Can We Trust Item Response Theory for AI Evaluation?

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multim
arXiv cs.AI 15d ago Healthcare

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains 44 clinician-reviewed
arXiv 15d ago Safety & alignmentHealthcare

The Industrialization of Research ; On AI-Driven Science and Its Consequences

Artificial intelligence is transforming scientific research -- not merely as a more powerful instrument, but as an autonomous participant in the research cycle itself. This transition constitutes, in the most precise sense of the term, the industrialization of research: a shift from a craft model, in which knowledge, method, and judgment are embedded in the researcher, to a pipeline model, in which these steps are decomposed, automated, and supervised. The US Department of Energy's Genesis Missi
arXiv cs.AI 15d ago Environment

Scaling Behavior Foundation Model for Humanoid Robots

Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to f
arXiv cs.AI 15d ago Agents & autonomyEnvironment

Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies

Online encyclopedias shape political opinion and, through it, democratic discourse. In late 2025, Grokipedia was released, an encyclopedia written entirely by the LLM Grok. One motivation behind the project was to provide an unbiased alternative to Wikipedia, which has faced accusations of "left-wing" and "liberal" bias. But does an encyclopedia written by an LLM deliver greater neutrality, or does it simply embed a different ideology? We conduct a large-scale political bias study on Grokipedia
arXiv 15d ago Bias & fairnessTransparency

Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents

AI coding agents set up projects by reading documentation and installing the dependencies it lists, without verifying their names, sources, or known vulnerabilities. By editing only a README, requirements file, or Makefile, an attacker can redirect the agent to an untrusted registry, a known-vulnerable version, or a wrong-but-plausible name: documentation becomes a vector for code execution. We present the first systematic evaluation of package-install-time supply-chain attacks delivered through
arXiv cs.HC 15d ago Military & securityAgents & autonomy
← Newer Older →