03:10 UTC

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

arXiv:2508.05775v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language generation and understanding. Meanwhile, they pose risks by inadvertently producing toxic, offensive, or biased content. This dual role of LLMs, both as powerful tools for text generation and as potential sources of harmful language, presents a pressing sociotechnical challenge. In this survey
arXiv cs.CY yesterday Bias & fairness

Measuring the State of Open Science in Transportation Using Large Language Models

arXiv:2601.14429v2 Announce Type: replace-cross Abstract: Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of their practice within transportation research remains under-investigated. Key features of open science, defined here as data and code availability, are difficult to extract due to the inherent complexity of the field. Previous work has either been limited to small-scale studies due to the labor-intensive n
arXiv cs.CY yesterday Jobs & economyFinance, VC & PE

FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets

As autonomous drone deployments scale from individual units to coordinated swarms, the human operator's role shifts from direct piloting to high-level supervision. Current interfaces often treat multi-drone control as a scaled-up version of single-drone operation. We instead investigate how reframing fleet supervision as spatial interaction can better support the spatial, temporal, and safety demands of complex missions. We present FleetScape, a Mixed Reality (MR) sandtable system that externali
arXiv cs.HC 2d ago Finance, VC & PE

From Micro-Cognition to Self-Construction: A Four-Layer Integrative Review of Psychological Theories in HCI

Human-computer interaction (HCI) is undergoing a paradigm shift from "tool use" toward "partnership" and even "mind symbiosis," with psychology evolving from a supplementary explanatory tool to a core pillar shaping interaction paradigms and long-term relationships. This paper systematically reviews relevant research and proposes a four-layer integrative framework comprising the Micro-cognitive, Meso-affective, Macro-social, and Self-constructive layers. The framework reveals that: the cognitive
arXiv 2d ago

Sensor-Placement-Agnostic Sonomyography: Toward Continuous High-Dimensional Control by Users with Tetraplegia

Sonomyography (SMG) enables continuous device control via ultrasound-measured muscle deformation signals, but existing SMG interfaces generally require substantial user- and sensor-location-specific training data and provide only one proportional signal or task-specific classification. We present a real-time, sensor-placement-agnostic SMG control system based on sparse optical flow tracking that enables continuous 1-DOF control after minimal calibration (3 pose definitions). We also present a pr
arXiv cs.HC 2d ago Privacy

Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment

Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi samples for every target class, limiting their ability to recognize unseen activities. We present Zero-Fi, a contrastive signal-language alignment framework for zero-shot Wi-Fi-based human activity recognition. Zero-Fi learns unified representations from complementary Wi-Fi signal features and aligns them with the semantic representations of nat
arXiv 2d ago Safety & alignment

(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding

Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewing may harm their understanding, impeding oversight, learning, and communication. To probe this, we have 54 students create a website with one of two AI systems: an agent that edits user code; or a chatbot where users write code alone or adapt generic code snippets. We test understanding via comprehension questions and a task where users extend t
arXiv cs.HC 2d ago Jobs & economyChildren & education

Information bubble and skill evolution: a theoretical framework for integrating distributed cognition, cognition offloading, and metacognitive regulation based on AIOE

Frontiers in Artificial Intelligence 2d ago Regulation

Learning faults in time: sequential behavioural modelling for complex fault detection in multi-robot systems

Reliable fault detection in multi-robot systems requires models capable of capturing complex, time-dependent fault signatures that manifest over extended temporal horizons rather than instantaneous observations alone. Existing data-driven approaches operate reactively on behavioural snapshots, failing to capture fault modes whose discriminative signature depends on temporally ordered precursors. This work formalises a theoretical impossibility result demonstrating that memoryless classifiers are
Frontiers in Robotics and AI 2d ago Bias & fairnessAgents & autonomy

Requiem Without an Orchestrator: A Commentary on Hollanek and Nowaczyk-Basińska (2024)

Hollanek and Nowaczyk-Basińska (2024) analyse the harms of AI-enabled re-creation services and offer four recommendations to providers. I accept their diagnosis but argue that their prescription shares a premise with the industry it seeks to reform: that a responsible provider remains present to carry it out. All four recommendations (retirement procedures, meaningful transparency, adult-only access, and mutual consent) are addressed to a continuing operator. Yet the paper’s own opening example,
Philosophy & Technology 2d ago HealthcareTransparency

What if I Want to be an Avatar: Moral and Legal Implications of Opting for a Life in the Metaverse

The rise of virtual worlds, collectively known as the Metaverse, compels a re-examination of longstanding debates concerning personal freedom, moral responsibility, and legal accountability. These immersive digital environments differ from prior communication technologies not merely in degree but in the phenomenological quality they produce: a pervasive sense of presence that blurs the boundary between the virtual and the physical. This article examines the ethical and legal challenges that aris
Science and Engineering Ethics 2d ago TransparencyEnvironment

DrugPred: an EdgeConv-GNN and Bio_ClinicalBERT based polypharmacy ADR prediction and specialist recommendation model

Adverse drug reactions (ADRs) are caused by medication and are considered a serious issue in healthcare when there is simultaneous use of different medications resulting in drug-drug interaction (DDI). Traditional approaches mostly focus on the effects caused by a single drug, and they fail to capture the side effects from drug combinations. In this research work, a deep learning-based DrugPred framework is proposed to predict ADR risks by integrating individual drug effects, interaction statist
Frontiers in Artificial Intelligence 2d ago Healthcare

Bridging agronomic science and context specific farm-level advisory through generative AI for rice systems in India

Agriculture is increasingly characterized by a data paradox, while the sector generates massive volumes of genomic, climatic, remote sensing, and field data. Translating this information into actionable, farm-level insights remains a critical bottleneck. Traditional advisory mechanisms cannot operate at the spatial scales or provide the context-specificity needed for climate adaptation and food security. The work presents GenAI as a transformative interface that makes advanced agricultural scien
Frontiers in Artificial Intelligence 2d ago EnvironmentBiotech

Continuous assurance for AI-driven clinical decision support systems

Healthcare systems are rapidly embedding adaptive and generative AI into core clinical processes. The integration of Artificial Intelligence into Clinical Decision Support Systems (AI-CDSS) highlights a fundamental transformation within healthcare delivery. This transformation enables advanced predictive analytics, multimodal data integration, and real-time augmentation of clinical decisions. However, AI introduces systemic, ethical, operational, and governance risks that challenge traditional h
Frontiers in Artificial Intelligence 2d ago RegulationHealthcare

A comparative evaluation of quantum machine learning architectures for breast cancer classification using clinical and genomic data

IntroductionIn recent years, high-dimensional clinical and genomic data have gained significant importance for prognosis and personalized medicine in breast cancer. But the use of quantum machine learning (QML) on such data is limited by the availability of few qubits, the computation time of quantum simulation, and dimensionality reduction. This work systematically compares several QML architectures for breast cancer classification in the presence of realistic and simulator constraints.MethodsT
Frontiers in Artificial Intelligence 2d ago HealthcareBiotech

A hierarchical federated learning framework with FedNova, game-theoretic matching, and QKD-assisted privacy for the internet of vehicles

The Internet of Vehicles (IoV) supports essential intelligent transportation applications but encounters challenges in federated learning (FL) due to non-independent and identically distributed (non-IID) data, vehicle mobility, resource heterogeneity, and strict privacy requirements in latency-sensitive scenarios such as misbehavior detection and accident response. Traditional FL methods, such as random client selection and standard FedAvg, often experience slow convergence and reduced performan
Frontiers in Artificial Intelligence 2d ago Privacy

“But it sounded confident”: the role of accuracy, tone, and disclaimers in users' medical decision-making

IntroductionArtificial intelligence (AI)-powered chatbots are increasingly used in healthcare for applications ranging from symptom triage to lifestyle guidance. Their effectiveness depends not only on their ability to provide reliable information but also on users engaging with their advice while remaining aware of potential inaccuracies. This study investigated how users perceive AI-generated medical advice, with a particular focus on the roles of accuracy, conversational tone, and disclaimers
Frontiers in Artificial Intelligence 2d ago HealthcareFinance, VC & PE

Overview of RAG-based and LLM-based approaches to personalization in healthcare AI applications

Recent advances in Large Language Models (LLMs), driven by transformer architectures such as Generative Pre-Trained Transformer (GPT), are opening new frontiers in healthcare Artificial Intelligence (AI) by enabling clinically relevant interactions between patients and clinicians. Yet persistent challenges—including limited real-time knowledge access, safety concerns and insufficient patient-centered contextualization—indicate that current systems often fall short in delivering efficient and rel
Frontiers in Artificial Intelligence 2d ago Healthcare

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. gener
arXiv cs.HC 2d ago Regulation

Incast-Free MoE Rate-Based Scheduling

Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks. In this paper, we demonstrate that RR causes a previously-undiscovered exponential incast phenomenon with MoE traffic. We propose an alternative proactive fair scheduling framework tailored for MoE workloads, which effectively prevents fabric oversubscription. We also outline how it can be implemented in NICs. Finally, through ext
arXiv fairness query 2d ago

Designing Needs- and Attention-Aware AI Learning Tools for Engineering Education: Insights from Psychological Outcomes

Artificial Intelligence (AI) is transforming higher education, but its benefits can vary depending on where, how, and how often it supports learning. While prior research emphasizes cognitive and academic outcomes, this study examines how AI chatbots support the psychological needs and motivational states of engineering students. A survey of college engineering students (n = 206) examined perceived effects of AI chatbots on autonomy, relatedness, and relief from competence frustration. Structura
arXiv cs.HC 2d ago Children & educationAgents & autonomy

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However, commonly used report-derived labels for pathology classification or generic image quality metrics for reconstruction may not reliably reflect clinical judgment. We systematically investigate how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA). To enable controll
arXiv cs.LG 2d ago HealthcareFinance, VC & PE

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninfo
arXiv 2d ago Healthcare

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents acro
arXiv red teaming query 2d ago Agents & autonomy

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduce AgentGUI, a user-friendly, locally hosted GUI for seamlessly observing and steering AI agents amid multiple concurrent, long-running sessions. AgentGUI features 1) rich agent trajectory visualizations, 2) effective manual and automated steering, an
arXiv cs.HC 2d ago Agents & autonomy

Weak-to-Strong On-Policy Distillation

On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs. Prevailing approaches assume a teacher at least as capable as the student: they either distill a larger model into a smaller one, which fails at the frontier where no larger teacher exists, or consolidate multiple domain experts trained from a shared base, which requires costly training at the student's
arXiv cs.LG 2d ago RegulationChildren & education

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for learning skills across tasks. SkillRise organizes related
HuggingFace Daily Papers 2d ago Agents & autonomy

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a pr
HuggingFace Daily Papers 2d ago Agents & autonomy

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V to L to A pathway as a direct V + L to A mapping. Instead of using a lar
HuggingFace Daily Papers 2d ago RegulationAgents & autonomy

The future of fact-checking in the algorithmic society

Fact-checking saw a rapid expansion in the mid 2010s when major social media platforms, especially Meta, started funding these activities. Today, however, fact checkers stand at a critical juncture. The early 2020s brought them three interrelated yet distinct crises: financial, technological, and legitimacy crises. The post The future of fact-checking in the algorithmic society first appeared on HKS Misinformation Review .
HKS Misinformation Review 2d ago Misinformation

User-Reported Misinformation Exposure Across Social Media Platforms

In this study, we surveyed users for their perception of misinformation exposure across social media platforms. Such perceived exposure is important because individuals' beliefs about how often they encounter false information can shape their trust in institutions, platforms, and even their friends. In a survey of 1,010 United States residents, we found that perceived exposure to misinformation varies substantially across platforms and is only moderately correlated with the frequency of platform
arXiv cs.HC 2d ago Misinformation

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems. It enables multiple agents to exchange arguments, critique each other's outputs, and iteratively converge towards a solution. However, research remains fragmented, with inconsistent terminology and no rigorous synthesis of MAD design dimensions. We present a systematic literature review characterizing 141 primary studies on MAD. We derive a three-dimensi
arXiv 2d ago Agents & autonomy

Development of a Blockchain-Based Platform to Enable Indigenous Data Sovereignty and Shared Research Participation With Indigenous Communities: Technology Prototyping and Community Engagement Study

Background: Historic and ongoing problematic practices regarding the collection, storage, and use of Indigenous health data have led to the need to ensure principles of Indigenous Data Sovereignty (IDS) are followed in research practices and technology development. Objective: This project, a partnership between UC San Diego and the Native BioData Consortium (NativeBio), sought to explore the practical application of blockchain technology and its potential to facilitate Indigenous-led research co
JMIR (Journal of Medical Internet Research) 2d ago Healthcare

On Exercising Governance Power in Decentralized Autonomous Organizations

A decentralized autonomous organization (DAO) is a governance entity that allows its stakeholders to manage blockchain-based protocols through smart contracts. The DAO explicitly specifies how stakeholders make and enforce decisions concerning a protocol's operation in a smart contract, aptly referred to as its governance contract. The design of this governance contract, therefore, has far-reaching implications for the security (trust) and privacy (transparency) of the smart contracts managed by
arXiv 2d ago RegulationPrivacy

Effectiveness and Implementation of Digital Health Interventions on Physiological, Psychological, and Functional Outcomes in Adults With Multimorbidity: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Multimorbidity involves heterogeneous disease combinations, treatment burden, competing priorities, and complex care pathways. Digital health interventions (DHIs) may support monitoring, self-management, and care coordination, but their effects on health-related outcomes remain uncertain. Objective: This systematic review and meta-analysis evaluated the effectiveness of DHIs on physiological, psychological, and functional outcomes in adults with multimorbidity, summarized implementat
JMIR (Journal of Medical Internet Research) 2d ago Healthcare

The Performance of ChatGPT-4o and DeepSeek-R1 in Interpreting Thyroid Nodule Ultrasound Text Reports: Multicenter Study

Background: Although thyroid nodules are detected in up to 60% of adults on ultrasound, the vast majority are benign, creating a substantial decision-making burden compounded by heterogeneous practice guidelines. Large language models (LLMs) show promise in processing unstructured medical text and are emerging as tools for report interpretation among both clinicians and patients. However, their reliability across distinct clinical tasks in thyroid ultrasound interpretation remains poorly charact
JMIR (Journal of Medical Internet Research) 2d ago Healthcare

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a goal, we should test whether lessons learned from one area transfer to the other areas. We study three such transfers, each taking a lesson developed in one SFT setting and testing it in another. First, we port a lesson about behavior generalization from alignment traini
arXiv cs.LG 2d ago Safety & alignment

Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation

Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as auxiliary model inputs. This limits model predictions to isomer identification rather than full molecular structure prediction. Although transformer models have been shown to identify molecular isomers with high accuracy, their reliability for unconstrained structure e
arXiv cs.LG 2d ago Safety & alignment

When benchmark inferences do not compose: Projectibility in AI evaluation

An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study
arXiv 2d ago

A Large Language Model–Driven System for Advance Care Planning Training Among Health Care Providers in the Chinese Context: Development and Technical Evaluation

Background: With the expanding need for advance care planning (ACP), innovative educational strategies for training health care providers are increasingly required. Large language model (LLM)–based ACP chatbots offer a novel and potentially effective solution to enhance health care providers’ competence in navigating complex ACP conversations. Objective: This study aimed to develop a Chinese-context ACP corpus to support an LLM-based chatbot and evaluate the feasibility and performance of a mult
JMIR (Journal of Medical Internet Research) 2d ago Healthcare
← Newer Older →