Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
HSBC UK trials AI shopping with Visa
HSBC UK is working with Visa to develop AI-powered shopping experiences, introducing secure technologies that will enable customers to use their cards in agentic commerce in a trusted and controlled way.
How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
As AI agents gain prevalance, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as hallucination, can cause agents to leak private information to third parties. As autonomous systems, agents also present the more active danger of performing sensitive tasks, such as bank transactions, without the user's intent or authorization. Recognizing this challenge, the agentic security community has developed numerous proposals for secure agentic systems. Mu
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
Agentic AI systems increasingly act through heterogeneous runtimes: local coding hooks, SDK tools, browser automation, managed-agent traces, API gateways, and workflow engines. A single operational act such as publishing code, changing identity state, moving money, or exporting data may therefore be represented by many incompatible runtime records. This makes a basic governance question difficult to answer: what action was actually approved, what evidence binds the approval to execution, and can
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely B
NVIDIA and Japan Bring Full-Stack AI and Robotics to Every Industry
Home to leading manufacturers, robotics pioneers, infrastructure builders and iconic gaming companies, of course, Japan is one of the world’s centers of AI — building across the full stack with NVIDIA technologies. This week NVIDIA and its partners in Japan are showcasing the AI ecosystem’s latest advancements. Check back here for updates.
Social Simulations: from Agent-Based Modeling to Digital Twins
This book chapter covers the evolution of social simulation from classical agent-based models, in which agents interact according to explicitly defined behavioral rules, to AI-enhanced simulations based on Large Language Models and, ultimately, Social Digital Twins: high-fidelity, data-driven representations of real-world socio-technical systems. Along this trajectory, we discuss the main methodological foundations, applications, advantages, and limitations of each paradigm, highlighting the pro
When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization strengthen or weaken? We study this question in open-source software, where bots open pull requests, review code, and merge changes alongside people, leaving a public record of every interaction. Treating bots as participants rather than tools, we examine 2,991 GitHub projects for two years before and after each adopted its first bot. We measure three capabi
Fraunhofer eröffnet drei Humanoid Robots Experience Labs in Deutschland
Das Fraunhofer IOSB schafft mit drei Humanoid Experience Labs ein niedrigschwelliges Angebot für Wirtschaft und öffentliche Hand zum Einsatz humanoider Roboter.
Explaining Reinforcement Learning Agents via Inductive Logic Programming
Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targeting the needs of a specific audience and lacking shared evaluation metrics. On the other hand, logic-based approaches within eXplainable Artificial Intelligence (XAI) provide compact, human-readable abstractions of decision-making. However, the syste
From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception
Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the integration of language understanding, environment perception, and autonomous navigation. This work presents a language-driven navigation framework that enables mobile robots to interpret user requests in natural language to move the robot to a destination and autonomousl
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode. This setting simplifies the problem and overlooks a more realistic requirement: an agent often needs to first find a language-described target and then persistently follow that target in a dynamic environment. While recent work has started to study human search, existing settings are typically evaluated in task-spec
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and still failed to act (the knowing-doing gap). Existing evaluations cannot separate these two failures; their reference policies either read privileged information the agent never sees, or are missing altogether. We introduce STOCKTAKE, a 26-week supply-
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System
Automatic scientific discovery has long been a goal of computational scholars - a machine that can discover nature's secrets on its own, moving computational systems beyond data-fitting tools toward the generation and refinement of mechanistic models of the universe. Recent advances in symbolic regression (SR) and large-language-model (LLM)-based agents suggest that such systems can recover equations from data, incorporate domain priors, and automate parts of the research workflow. However, most
Legatics’ New MCP Server Connects Your AI Tools
Legatics, the well-known transaction management platform, has launched a Model Context Protocol (MCP) server. This will allow your ‘AI assistants or agents to connect directly ...
Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis
Systematic comparisons between current situations and structurally similar past events in the historical, i.e., historical analogies, is among the most powerful tools for foresight analysis. In this work, we present a new task called Analogical Deep Research (ADR) to Large Language Model (LLM) agents and construct the first ADR benchmark ADR-bench to study whether LLM agents are able to find and leverage historical analogies when doing foresight analysis. Our investigation reveals a key obstacle
Semantic Anchoring for Robotic Action Representations
Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. A fundamental question therefore arises: what constitutes a good action representation? Inspired by the mirror neuron theory's insight that observation and execution share an intention-level encoding, we examine whether a robot's action representations preserve the semantic structur
SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a guard model that labels each proposed action as safe or unsafe, but this binary view conflates two distinct decisions: whether the action is harmful in itself, and whether it is appropriate given the user's context. It also operates at the granularity of action categories rather than individual instances, producing routine interruptions that erode a
News June 30, 2026 New CLTC White Paper Introduces Method of Evaluating Privacy and Security of AI Agents
AI Coding Agents Fail at Teamwork
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is
OpenAI's Codex now encrypts instructions between AI agents, leaving developers blind to internal delegation
Since early June, OpenAI's coding tool Codex encrypts the instructions a main agent passes to its subagents. Developers can no longer track how tasks get delegated internally. For the larger GPT-5.6 variants Sol and Terra, the encryption is mandatory. The article OpenAI's Codex now encrypts instructions between AI agents, leaving developers blind to internal delegation appeared first on The Decoder .
Agile perceptive multi-skill locomotion for quadrupedal robots in the wild
Enabling quadrupedal robots to traverse complex terrains-from rugged outdoor environments to urban landscapes-requires seamless integration of multiple motor skills, smooth transitions between gaits, and high-speed perceptive locomotion using only onboard sensors. We present APT-RL (Action Pretrained Transformer-based Reinforcement Learning), a unified framework that enables multi-skill locomotion to achieve high-speed traversal in complex environments through autonomous skill transitions utiliz
AI Startups To Watch: 5 Startups That Caught Our Eye In July
The Indian AI ecosystem is no longer restricted to chatbots, agents and copilots. Across industries, tech is increasingly being deployed…
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come from a phone, be processed on a desktop, and the result may need to appear on another device. Most existing benchmarks center on a single dominant execution environment, making it difficult to evaluate whether agents can acquire and integrate information across heterogene
When Artificial Intelligence Breaks Our Moral Vocabulary: Structural Stalemate, Conceptual Insufficiency—and Why Human Rights Can’t Do the Work Alone (But We Can’t Do ...
It has become a familiar refrain: we are not conceptually prepared for artificial intelligence (AI). We throw around words like intelligence, understanding, autonomy, and consciousness as if their ...
AgentSociety 2: An Integrated Research Environment for Executable Social Science
arXiv:2607.11895v1 Announce Type: new Abstract: AI scientist systems are beginning to automate parts of scientific research, but social science poses a distinct challenge: its objects of inquiry are not merely datasets or laboratory protocols, but integrated social processes involving situated participants, interaction contexts, interventions, and outcomes. Yet a critical link is missing: existing systems either assist isolated research tasks or simulate agents as experimental subjects, leaving
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
arXiv:2607.11999v1 Announce Type: new Abstract: From maritime trade to commercial nuclear power, insurance has been the enabler of major economic and technological developments by pricing risk, limiting downside, and spreading best practices. The emerging AI agent economy, projected to handle trillions of dollars in transactions by 2030, looks to be the next such development. Yet insurers' exposure to AI agent risk currently sits largely unpriced across existing insurance lines; between this sil
CityBehavEx: A Scalable and Empirically Validated LLM-Assisted Urban Simulation Platform
arXiv:2607.12086v1 Announce Type: cross Abstract: Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validated against empirical mobility patterns. We present CityBehavEx, an interactive LLM-assisted urban simulation platform that scales to city-size populations, exposes agent behavior for inspection, supports empirical validation, and generates mobility patterns that better match real-world spatial, te
How Agentic Is Agentic Commerce? A Population-Scale Measurement of x402 Adoption and Authenticity
arXiv:2607.12575v1 Announce Type: cross Abstract: AI agents are said to be forming an economy in which they pay, on their own, for the data, APIs, and compute they consume. x402, which settles a stablecoin payment on-chain for each purchase, is the most widely deployed protocol for this, and its hundreds of millions of settlements are read as proof that the economy has arrived. We show the count cannot be read as adoption: it is the one metric an interested party can manufacture almost for free,
Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs
arXiv:2607.12650v1 Announce Type: cross Abstract: Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Verified Agentic Reasoning), a Lean 4-based tool-calling architecture in which the Lean kernel is the sole minter of Verified claims via tool-attestation axioms and declared source lifts. Every verified output structurally
Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems
arXiv:2607.12755v1 Announce Type: cross Abstract: AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" with respect to core functions and capabilities. I.e., many systems take inputs and give outputs, but without users having any ability to see how the former lead to the latter. AI-enabled systems are also being used to augment autonomy in systems, and autonomy coupled with opacity raises numerous concerns surrounding, e.g., the reliability o
Xiaomi updates progress on humanoid robots in auto factory, achieves 98% success rate in some tasks
Xiaomi has revealed the latest progress of its humanoid robot deployed in an automotive factory. Following four months of iteration, the robot’s success rate at a self-tapping nut loading station has risen from 90.2% to 98%, narrowing the gap with human workers’ qualification rate to just one percentage point, according to the company. The company […]
Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledge Component (KC) models across integrated Chat, Tutor, and Quiz modes. That prior work validated LEA on a single STEM course (CMP511) exclusively through simulation, using synthetic learner agents. This paper extends that work by reporting the first
TANDE: Disentangling Verbal and Nonverbal Backchannels in Emotional AI-Avatar Conversations with Young Adults
Embodied conversational agents (ECAs) need effective empathic grounding to foster social support and engagement. Expanding into emotional domains, ECAs now use Large Language Models (LLMs) and multimodal human-agent interactions to enhance their capabilities. Yet, understanding the impact of backchanneling modalities on young adults and their gender remains limited. We introduce TANDE, an LLM-powered ECA designed for emotional conversations with young adults, a population experiencing mental, pe
Cribl Adds Agentic Detection Engineering & Boosts SecOps With CardinalOps Deal
CardinalOps will give Cribl customers the ability to map detection rules and security controls to the MITRE ATT&CK framework. SecOps teams can identify coverage gaps and operationalize threat intelligence.
Benchmark evaluation in task and motion planning using iteratively deepened AND/OR graph networks
In robotics research, each subdomain presents a distinct set of challenges, and any framework designed for a given domain must effectively address these complexities. However, a single application within that domain may not fully capture the breadth of challenges inherent to it. To enable systematic and comprehensive evaluation, the robotics community has developed standardized problem scenarios and associated performance metrics, commonly referred to as benchmarks, which collectively represent
Object-centric diffusion policies for real-world robotic-arm imitation learning
Imitation learning in complex, unstructured environments remains challenging due to the difficulty of grounding perception in physically meaningful representations and the need to model multimodal action distributions. Existing approaches often rely on unstructured pixel-level feature encodings or stochastic latent-variable decoders, which can lead to brittle attention in cluttered scenes. In this work, we present a novel integration of detector-based visual representations with conditional diff
Human-like conversational agents as social partners: a scoping review of socioaffective mechanisms, well-being outcomes, risks and governance in the post-Turing era
IntroductionLarge language models have evolved from laboratory demonstrations into mass-market companion-style conversational agents that many users treat as social partners. As these systems produce increasingly human-like conversational behavior, users may attribute mind, form affective bonds, disclose sensitive information, and rely on agents for emotional support, creating both potential benefits and psychosocial risks.MethodsWe conducted a PRISMA-ScR-informed scoping review of socioaffectiv
5 Trends That Defined AI Engineering at World’s Fair 2026
At this year's AIE World’s Fair, AI engineering entered a new phase: building systems around agents, rather than just building with agents.
Man Killed by Vehicle While Fleeing ‘Encounter’ With Federal Immigration Officers in Florida, Official Says
This marks the third death in a week involving encounters with federal immigration agents.