17:00 UTC

An offline approach to fNIRS-guided reinforcement learning for robot behavior

Human-in-the-loop Reinforcement Learning has become a popular approach to training, finetuning, and aligning robot behavior with user preferences. Our paper explores the feasibility of using brain signals via functional near-infrared spectroscopy (fNIRS) to modulate robot learning in simulation. We compare agents trained on passive (observational) versus active (demonstrative) interaction tasks, and test multiple methods for enhancing the RL algorithm with the neural signal, focusing on paramete
arXiv 16d ago Agents & autonomy

Developing a Core Outcome Set for the Evaluation of Remote Patient Monitoring Interventions Using the Sextuple Aim: Modified Delphi Study

Background: The rapid expansion of remote patient monitoring (RPM) interventions highlights the need for their comparison and evaluation. Current evaluation frameworks often fail to capture a multistakeholder perspective. Traditional health technology assessment approaches emphasize health and economic outcomes, providing an incomplete picture of RPM’s broader impact. The Sextuple Aim, encompassing health outcomes, costs, patient and provider experience, equity, and sustainability, offers a more
JMIR (Journal of Medical Internet Research) 16d ago Bias & fairnessHealthcare

Federal Cuts and Public Health: Social Media Sentiment Among Federal Employees

Using sentiment analysis of 44,216 public health–related posts from the FedNews Reddit forum (2020-2025), we found that fear and anger scores rose 32% and 63%, respectively, in 2025 over 2024. These results underscore the adverse mental health impact of federal workforce changes and funding cuts on public health employees.
JMIR (Journal of Medical Internet Research) 16d ago Jobs & economyHealthcare

Dysco: Dynamic Subspace Boosting to Mitigate LoRA Interference in Federated Learning

Federated fine-tuning of large pre-trained models increasingly relies on Low-Rank Adaptation (LoRA) to reduce communication and computation, but heterogeneous clients can make adapter aggregation unstable. We identify the data-parameter interference as a geometric source of this instability. This interference is controlled by the alignment between LoRA update subspaces and client activations, suggesting that federated LoRA aggregation should be viewed not only as parameter averaging but also as
arXiv cs.LG 16d ago Safety & alignment

Perceived Importance of Abortion Care Features and Access to Telehealth Technologies Among Medication Abortion Patients by Abortion Care Model: Cross-Sectional Analysis of a Prospective Cohort Study

Background: Medication abortion accounts for the majority of abortions in the United States, driven in part by the growth in access to telehealth provision of medication abortion. While research indicates high patient satisfaction with telehealth medication abortion care, research on care preferences of patients who use medication abortion remains understudied. Understanding these preferences is essential to informing evidence-based policies that enable people to access person-centered abortion
JMIR (Journal of Medical Internet Research) 16d ago Healthcare

When Is Delegated Play Truthful? Within-Range Regret and the Trilemma of Aligned Delegation

Advertisers delegate bidding to autobidders; users delegate tasks to language-model agents. A person describes what they want to an automated proxy that acts in a mechanism on their behalf. This is the revelation principle in production, and it forces a question classical theory assumes away: when is it optimal to describe yourself honestly to your own proxy? We show the answer turns on one quantity, the proxy's within-range regret. The most a principal can gain by misreporting equals the regret
arXiv red teaming query 16d ago Agents & autonomy

Application of AI in Hypertension Health Education: Scoping Review

Background: Hypertension is a major global health challenge, and effective health education is crucial for improving patients’ self-management. Traditional health education approaches are often limited by insufficient personalization, accessibility, and scalability. Artificial intelligence (AI), including natural language processing, machine learning, and large language models (LLMs), offers promising solutions to address these limitations. However, evidence regarding AI applications in hyperten
JMIR (Journal of Medical Internet Research) 16d ago HealthcareChildren & education

Effects and User-Reported Experiences of a Self-Management Mobile Health App for Grieving Adolescents: Randomized Controlled Trial

Background: Adolescents who experience the loss of a family member are at increased risk of adverse mental health outcomes, yet many face barriers or may be reluctant to access in-person or group-based support. mHealth (mobile health) interventions can help address these barriers by offering flexible, accessible, and low-threshold support. Objective: This study evaluated the short- and long-term mental health effects of Alba – Youth in Grief, a preventive self-management mobile app for bereaved
JMIR (Journal of Medical Internet Research) 16d ago Healthcare

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI.
arXiv 16d ago Finance, VC & PE

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI.
arXiv cs.LG 16d ago Finance, VC & PE

Comparative Analysis of Expert, Clinician, and Health Care User Interactions With Summary of Findings Tables: Usability Study

Background: Summary of findings (SoF) tables are widely used in systematic reviews and clinical practice guidelines to present evidence about health care interventions in a concise and transparent format. Although developed to improve accessibility and interpretation of evidence, previous studies have shown that users often experience difficulties understanding statistical information, certainty ratings, and the relationships between outcomes and treatment effects. Limited research has explored
JMIR (Journal of Medical Internet Research) 16d ago HealthcareTransparency

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can become trapped in repetitive loops, wasting search budgets and ultimately compromising the quality and completeness of the final output. We introduce SearchOS, a system-level multi-age
HuggingFace Daily Papers 16d ago Agents & autonomy

BadWAM: When World-Action Models Dream Right but Act Wrong

World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluatin
HuggingFace Daily Papers 16d ago Safety & alignmentAgents & autonomy

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between episode-level outcomes and token-level policy learning. We propose SEED (SElf-Evolving On-Policy Distillation), a self-evolving
HuggingFace Daily Papers 16d ago RegulationAgents & autonomy

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturb
HuggingFace Daily Papers 16d ago RegulationAgents & autonomy

Video = World + Event Stream

We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and
HuggingFace Daily Papers 16d ago Environment

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to CAD models, yet typically their correspondences are appearance-driven and degrade under occlusion or sim-to-real domain shift. To address these limitations, we introduce SUFLECA (Scaling Up Feature LEarning for CAD Alignment), a we
HuggingFace Daily Papers 16d ago Safety & alignmentAgents & autonomy

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment. The gap is especially important for AI agents, whose observations, tool outputs, documents, and prior decisions accumulate over long trajectories. LongStraw is an architecture-aware execution stack for million-token RL post-training under a fixed GPU bu
HuggingFace Daily Papers 16d ago Agents & autonomy

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In particular, the law of total probability asserts that prior-weighted conditional distributions aggregate into population-level marginals over any valid partition of the population. In t
HuggingFace Daily Papers 16d ago Regulation

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by
HuggingFace Daily Papers 16d ago Agents & autonomyEnvironment

On-Policy Delta Distillation

On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output distribution. The delta signal is
HuggingFace Daily Papers 16d ago Regulation

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial videos and other multimodal human resources largely underused. We present RESOURCE2SKILL, a framework that distills multimodal resources, including tutorial videos, repositories, articles, and reference artifacts, into executable skills for software agents. RES
HuggingFace Daily Papers 16d ago Agents & autonomy

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreemen
HuggingFace Daily Papers 16d ago Regulation

Multi-Turn On-Policy Distillation with Prefix Replay

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed pr
HuggingFace Daily Papers 16d ago RegulationChildren & education

PReM: Learning What to Preserve and When to Refresh for Context Compression

Efficient long-context inference is not only about reducing memory cost, but also about keeping useful contextual evidence accessible as generation proceeds. However, existing compression-oriented approaches, such as key-value (KV) cache compression and context compression, often either make an early decision about which contextual information to keep or rely on an external compressor. Such designs make it difficult to adapt the compressed context to the evidence needed by later reasoning steps.
arXiv 16d ago

The Relation Between eHealth Literacy and Online Health Information–Seeking Behavior: Systematic Review and Meta-Analysis

Background: Online health information–seeking (OHIS) behavior shapes health self-management, and eHealth literacy—the ability to seek, appraise, and apply electronic health information—is regarded as its key driver. Previous reviews aggregated heterogeneous outcomes, focused on measurement properties, or examined single clinical populations, without isolating the eHealth literacy–OHIS link. Objective: This study quantified the strength and heterogeneity of the eHealth literacy–OHIS association a
JMIR (Journal of Medical Internet Research) 16d ago Healthcare

Correction: Rising to the Challenge of Early Screening in Primary Health Care Through the Web Italian Network for Autism Spectrum Disorder (Win4ASD) in the Pediatric Population: Retrospective Observational Study

JMIR (Journal of Medical Internet Research) 16d ago Healthcare

NeuroGRIP: Retrieval-Augmented Graph Refinement for Knowledge-Grounded EEG Seizure Diagnosis

Seizure diagnosis from EEG signals is a critical yet persistently challenging task, due to the complicated neural dynamics and the spurious connections in inter-channel modeling. While spatial-temporal graph neural networks (STGNNs) have advanced EEG brain network representation learning, the resulting graph structures suffer from low clinical plausibility and limited interpretability due to their purely data-driven nature. To this end, we introduce NeuroGRIP, a retrieval-augmented graph refinem
arXiv cs.LG 16d ago Safety & alignmentHealthcare

Traccia: An OpenTelemetry-Based Governance Platform for AI Systems

The rapid development of Large Language Models (LLMs) and Artificial Intelligent (AI) powered autonomous agents has fundamentally changed the existing forms of software governance. In spite of the rigorous standards of transparency and account ability required according to the international frameworks such as the European Union's AI Act, there is a considerable gap between theory and reality. The present study discusses the inherent drawbacks of currently utilized platforms for LLM evaluation, m
arXiv 16d ago RegulationAgents & autonomy

Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)

As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, and educational equity. Existing measures of reliance were developed inductively, focused on discrete problem-solving tasks, and validated mainly with homogeneous samples. This study developed and validated the GenAI Reliance Types Scale (GenAI-RTS), a 20-item instrumen
arXiv cs.HC 16d ago Bias & fairnessChildren & education

ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs

Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict? We investigate this question in the context of tool-calling LLM agents deployed in regulated industries, where agents processing confidential documents may encounter content that triggers safety-trained values (e.g., public welfare) that conflict with deployment-context instructions (e.g., internal logging). To empirically verify this phenomenon, we build a benchmark of 128 scena
arXiv 16d ago RegulationSafety & alignment

Behavior Change Content and Implementation of Large Language Model–Driven Conversational Agents in Cardiometabolic Care: Scoping Review

Background: Large language models (LLMs) are increasingly embedded in conversational agents for cardiometabolic care. These systems could support self-management, but their behavior change content, delivery mechanisms, and implementation transparency are poorly understood. Objective: This scoping review mapped behavior change techniques (BCTs) used in LLM-driven conversational agents for cardiometabolic prevention and management, described how these techniques are delivered across static, rule-b
JMIR (Journal of Medical Internet Research) 16d ago Agents & autonomyTransparency

AI Agents Do Not Fail Alone:The Context Fails First

Context engineering has become central to building reliable AI agents, yet it remains largely unmeasured. Agents do not fail in isolation: their behavior is shaped by the instructions, tools, memory, retrieved knowledge, guardrails, and untrusted inputs accumulated in their context. When this context is weak, agents drift, hallucinate, misuse tools, ignore constraints, become vulnerable to injection, and waste tokens. This paper validates context-engineering quality as an independent leading ind
arXiv 16d ago Agents & autonomy

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation

Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases. Traditional active learning and manual annotation fail to scale against the complexity and volume of novel multimodal threats. In this paper, we propose an automated, agentic red-teaming framework that systematically synthesizes difficult examples using an iterative strategy that proposes novel hy
arXiv red teaming query 16d ago Safety & alignmentAgents & autonomy

Align AI to Dynamic Human-AI Workflows

Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions. In this paper, we argue for a shift from static and emulative to interactive and complementary alignment, where preferences emerge through interaction and alignment is defined not by satisfying preferences alone. We first formalize this gap by contrasting existing alignment with a
arXiv 16d ago Safety & alignment

Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection

Pretrained vision-language-action (VLA) policies provide strong language-conditioned manipulation knowledge, but they remain largely vision-driven and can struggle once manipulation enters contact states where the scene is occluded, depth is ambiguous, or small force errors push execution off the offline demonstration distribution. We present LIFT (Late Reactive Injection of Force for VLA Post-Training), a force-aware post-training framework that adds contact reactivity to a pretrained VLA polic
arXiv 16d ago

LIGO-PINN: Learned Initialization via Gated Optimization to Alleviate Convergence Failures in Physics Informed Neural Networks

Physics-informed neural networks (PINNs) have had a broad research impact in modeling domains governed by partial differential equations (PDE). However, PINNs have been shown to perform poorly, sometimes even converging to trivial solutions, in challenging PDE domains, or when generalizing to unseen but related PDE domains. Previously proposed solutions detail hyperparameter tuning to reduce loss imbalance between data-driven and physics guided losses, curriculum learning based training strategi
arXiv 16d ago

SeeSE3: Emergence of 3D Space in Vision Features

In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that probe 3D awareness of vision features by regressing image-centric quantities such as depth or normals, we investigate the relation between the structure of the space of visual features and the group of Euclidean transformations $SE(3)$. We propose a set of probes to evaluate this relation from both topological and geometric persp
arXiv 16d ago Finance, VC & PE

Privacy Leakage in Federated Learning in Radiology Reports: A Comparative Evaluation of Tokenizer-Driven Privacy Risks

Federated learning (FL) enables multi-institutional training on clinical text without sharing raw data, but gradient inversion can reconstruct sensitive information from shared model updates. The extent of this leakage for radiology reports, and the role of tokenizer design, remains unclear. We quantify gradient-based text reconstruction in FL and compare privacy risk across three tokenizers with the model architecture held fixed. Six FL clients trained a GPT-2-style transformer (sequence length
arXiv cs.LG 16d ago PrivacyHealthcare

Why I Left Google DeepMind

Alignment Forum 16d ago
← Newer Older →