Research · RSS feed
New papers on fairness, safety, alignment and governance.
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored. When no scratchpad is visible, behavior alone cannot distinguish strategic from genuine compliance. We ask whether hidden states reveal what outputs hide. We run a 13-model sweep for naturally-emerging faking, then probe and steer hidden states on the two models that fake. Natural faking appears only in Qwen3-32B (+18.2pp) and Llama-3.1-8B (+24.4pp at n=
Marker-free deformable registration and fusion for augmented reality-guided positive margin localization during tumor resection surgery
Positive margins in head and neck oncologic surgery require mapping specimen-side pathology findings to the patient resection bed. This is challenging because pathologists identify the positive margin on slices of the resected, deformed specimen, while surgeons must relocate the corresponding site on the resection bed using only verbal descriptions and no visual guidance. We present a marker-free augmented reality (AR) workflow for mapping a margin label from a three-dimensional specimen scan to
Benchmark evaluation in task and motion planning using iteratively deepened AND/OR graph networks
In robotics research, each subdomain presents a distinct set of challenges, and any framework designed for a given domain must effectively address these complexities. However, a single application within that domain may not fully capture the breadth of challenges inherent to it. To enable systematic and comprehensive evaluation, the robotics community has developed standardized problem scenarios and associated performance metrics, commonly referred to as benchmarks, which collectively represent
All too perfect: bias and aspiration in persona generation with LLMs
Synthetic data generated by large language models plays a central role in the training and alignment process of other AI systems. However, this process also risks inheriting the structural biases of organic corpora and embedding new biases that stem from the design choices underlying the data creation process. This paper examines the systematic biases that emerge when large language models (LLMs) are tasked with generating synthetic personas. We introduce a reproducible, minimally conditioned pi
Towards transparent financial AI: a systematic review of graph learning and explainable methods for credit risk and fraud detection
Graph-based learning and explainable artificial intelligence (XAI) are increasingly used to improve both predictive performance and transparency in financial risk modelling. This paper presents a systematic literature review of AI and machine learning approaches for credit risk assessment and fraud detection, with specific attention to graph-based methods and explainable frameworks. Following a PRISMA-guided methodology, 149 studies published between 2015 and 2025 were analysed across multiple a
Epistemic decolonisation and the integration of African traditional medicine: toward an epistemic redress model for inclusive AI in mental healthcare
Artificial intelligence (AI), particularly machine learning (ML) systems, is increasingly used in mental healthcare to support diagnosis and treatment. These systems analyse large datasets—including speech patterns, social media activity, and biometric indicators—to detect conditions such as depression and anxiety. Although often framed as objective and clinically neutral, these technologies embed epistemic and cultural assumptions derived largely from Global North contexts. This paper advances
Object-centric diffusion policies for real-world robotic-arm imitation learning
Imitation learning in complex, unstructured environments remains challenging due to the difficulty of grounding perception in physically meaningful representations and the need to model multimodal action distributions. Existing approaches often rely on unstructured pixel-level feature encodings or stochastic latent-variable decoders, which can lead to brittle attention in cluttered scenes. In this work, we present a novel integration of detector-based visual representations with conditional diff
Human-like conversational agents as social partners: a scoping review of socioaffective mechanisms, well-being outcomes, risks and governance in the post-Turing era
IntroductionLarge language models have evolved from laboratory demonstrations into mass-market companion-style conversational agents that many users treat as social partners. As these systems produce increasingly human-like conversational behavior, users may attribute mind, form affective bonds, disclose sensitive information, and rely on agents for emotional support, creating both potential benefits and psychosocial risks.MethodsWe conducted a PRISMA-ScR-informed scoping review of socioaffectiv
Artificial intelligence in cardiology: implications for healthcare outcomes
Artificial Intelligence (AI) has the potential to revolutionize medicine, particularly in the field of cardiology. There are significant diagnostics and treatments variabilities in the field of cardiovascular medicine that affects racial and ethnic racially and ethnically diverse populations as well as female patients across all age groups. The efforts put forth towards the development of AI and precision medicine within the cardiovascular practice do not fully account for existing variations in
OMNIS: a spatially informed multi-omics deep-learning framework for tumor recurrence prediction and primary–metastatic tumor differentiation title page
BackgroundCancer recurrence and distant metastasis are major causes of cancer-related death, yet existing biomarkers and single-omics models have limited accuracy and interpretability across tumor types.MethodsWe developed OMNIS (OMics Network Integration and Spatial representation), a convolutional deep-learning framework that embeds multi-omics profiles into a five-channel genomic image ordered by Hi-C–derived chromosomal proximity. Somatic mutation, copy-number alteration, DNA methylation and
Privacy Preserving Recommender Systems Balancing Personalization with Privacy
Personalized recommendation systems are central to modern e-commerce and retail platforms, but they typically rely on centralized storage of detailed user interaction data, creating significant privacy and regulatory challenges. With increasing requirements from regulations such as GDPR, CCPA, and CPRA, organizations must develop recommendation systems that preserve user privacy without substantially degrading recommendation quality. This work presents and evaluates a privacy-preserving recommen
Discourse-Aware Policy Analysis with Argumentation: A Hybrid LLM-Symbolic Framework for Disaster Governance
Policy documents shape governance outcomes, but their reasoning is often implicit. Participatory commitments and managerial control routinely coexist in the same text, and the tensions between them are rarely stated directly. Existing computational approaches to policy discourse cannot express the frame-mediated relations that drive these tensions, where one argument narrows or instrumentalizes another rather than rejecting it. End-to-end summarization by large language models produces fluent te
Guidance for Building Hospital at Home: Qualitative Descriptive Thematic Analysis of a Pan-Canadian Community Participatory Workshop Series
Background: Virtual health care models, such as (HaH), are an alternative to in-person care and allow hospitals to expand care capacity and delivery without the need for additional brick-and-mortar structures. While generally well received, there is an overall lack of awareness among those receiving and giving care about what HaH is and what it does, and uncertainty about the conditions needed to implement HaH in a safe, sustainable, and equitable way. Objective: In this descriptive qualitative
User Acceptability and Adoption of AI-Generated Lifestyle Intervention Recommendations: Scoping Review and Theoretical Integration
Background: Artificial intelligence (AI)–generated lifestyle recommendations are increasingly used to support health behavior change. However, AI advice does not necessarily mean that users will accept or adopt those recommendations. Although prior reviews have examined AI-enabled lifestyle interventions and health behavior technologies, fewer have focused on whether users accept and adopt AI-generated recommendations. Objective: This scoping review aimed to map user acceptability and adoption o
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We pre
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly generate future videos at inference time, incurring substantial computational overhead and hindering real-time closed-loop deployment. GigaWorld-Policy addresses this issue with an action-centered formulation, where future visual d
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device ecosystems and prevent performance improvements through continuous learning from execution experience. To resolve these issues, we propose the Know Deeply, Act Perfectly paradigm for personal assistants, which holds that accumulated user interaction and task-run
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Mod
Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely B
Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving
Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based testing, scenarios are defined as executable scripts. Yet automatically generating such scripts from regulatory descriptions remains an open challenge, and existing approaches face fundamental trade-offs. Retrieval-assemble methods achieve reasonable compilation rates but lack scalability, whereas retrieval-based full-script generation suffers from low compilation success rates. We pr
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagina
Cura 1T: Specialized Model for Agentic Healthcare
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated sel
Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings
Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if k verifier calls all accept it. Under conditionally independent gates, the recent Odds Law (arXiv:2606.15712) shows that posterior log-odds grow linearly in k, so failure decays exponentially, and states that "a tight theory of partially correlated verifier cascades remains open." This note gives a minimal such theory. Modeling the per-instance false-accept rate on the generator's
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in na
Diagnosing and Calibrating Tool-Call Boundary Drift in Multi-Teacher On-Policy Distillation
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool calls, another in direct responses, and the student can learn from both on its own generated distribution. We show that this strategy can induce a behavior shift that is invisible from aggregate losses alone. In a two-teacher tool-use setting, vanilla generalized knowledg
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support visual and semantic generalization. Co-training on web image-text data, a common remedy, does not prevent this; it applies language and action losses to separate observations, leaving VLAs with language-action misalignment that stan
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-token prediction. At inference time, it is preferred that an agent can decide whether to proceed with its current reasoning, defer to a stronger model, request additional information, invoke external tools, or abstain under the given setup. Existing approaches address these decisions through prompt-level routing, external orchestration, or task-specific fine-tuning, which primarily re
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interact with third-party services. This paper develops an AI-native mathematical framework for underwriting, pricing, and contract design for agentic AI deployments. A deployment is represented by a risk state that captures autonomy level, operational authority, permission exposure, governance maturity, and dependency concentration. The framework maps
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
Most AI-for-science systems focus on scaling a single reasoning process through better models, larger context windows, long-horizon agentic execution, or digital co-scientists working with one principal user. However, challenging scientific problems are rarely solved by one reasoner alone. They are solved by teams whose members bring different priors, experimental backgrounds, tacit knowledge, and domain-trained intuitions. The open problem is therefore not only how to scale models, but how to c
Patient- and Caregiver-Informed Considerations for the Design and Implementation of Generative AI–Supported Patient-Centered Clinical Decision Support: Qualitative Study
Background: Generative artificial intelligence (AI) has the potential to impact health care by transforming workflows and improving outcomes. Patient-centered clinical decision support (PC CDS) are digital tools that use patient-specific information and patient-centered outcomes research to improve health care decision-making. Generative AI is increasingly being incorporated into PC CDS tools. As patient-facing digital tools continue to expand within the health ecosystem, it is important to gath
Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference
Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (delimiters or whitespace) carries an order of magnitude more energy than any content role, and structural KEY to
SoftBoard: A Multi-Agent Tool for the Creation and Evaluation of Low-Fidelity Prototypes
User Experience (UX) is recognized as a critical factor for the success of digital products, particularly in software startups, environments marked by time constraints, limited resources, and low maturity in design practices. Building Minimum Viable Products (MVPs) through low-fidelity prototyping represents a well-established strategy for rapid validation cycles at reduced cost. A systematic literature mapping, however, revealed gaps in the ecosystem of available tools: a predominance of genera
Composable Trust for Language Models: A proven boundary and a measured defense
In a language model, instructions and data share one token stream, so nothing inside the model's generation can keep untrusted text from steering it. We develop a trust model that places the authority to act outside the model, in code: a source's standing, not its content, decides which operation runs and whether it acts. A lower-trust source may inform an answer but not override a higher one. An unmodified model runs inside a deterministic pipeline that ranks inputs by source integrity, and a f
Cultural, Organisational, and Individual Factors Contributing to Cyber Incident Reporting: A Systematic Literature Review
Publication date: Available online 14 July 2026 Source: Computers in Human Behavior Author(s): Rick van der Kleij, Olivier Spinnler, Julia Broderick-Hale, Katie Hendriks, Anthonie Drenth, Joshua van Wijgerden
Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution
Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-context-first strategy--re-reading files and dependencies they have already seen--turning a one-line edit into a small code-base audit. We argue the missing capability is task-aware execution-scope estimation: judging a task's difficulty, the information it truly needs, and the shortest reliable path be
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale
Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a procedural driving simulator and self-play training stack. A configurable C engine runs simulation on the CPU and policy inference on the GPU over a zero-copy path, sustaining 1.3M agent-steps per se
PalmClaw: A Native On-Device Agent Framework for Mobile Phones
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user
Dynamic Resource Allocation for Ensemble Determinization MCTS
Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elements of randomness and hidden information. In particular, several Monte Carlo Tree Search (MCTS) variants are commonly used in such domains. In this paper, we propose a series of enhancements for Ensemble Determinization MCTS, introducing two axes for dynamic resource allocation. First, Dynamic Number of Determinizations, increases or decreases the number of cu
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models
The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models. In this paper, we use a crowdsourced study to show that providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model, even when the model is patently discriminatory and unfair. Our results suggest that XAI designers and developers n