06:01 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL

Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predicting and pre-executing an agent's next tool call if the prediction matches the agent's eventual tool call, but existing speculators are typically separate draft models or cached traces that are poorly aligned with the deployed agent's own behavior. We identify this speculator-agent gap and show that the target agent itself is a strong next-call s
arXiv 2d ago Research Agents & autonomy

Wissenschaftler testen Ernteroboter auf Obstplantage am Bodensee

Am Bodensee wird ein Roboter getestet, der für den Einsatz im Obstanbau entwickelt wurde. Er soll künftig bei der Apfelernte eingesetzt werden.
Heise Online (DE) 2d ago News Agents & autonomy

GSA inks agentic AI OneGov deal with CORAS

CORAS is a FedRAMP High-certified agentic-AI “decision maker” that has already been authorized for use at the Department of Defense. The post GSA inks agentic AI OneGov deal with CORAS appeared first on FedScoop .
FedScoop 2d ago News Military & securityAgents & autonomy

Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Critic Controller

This paper presents a cooperative indoor UAV guidance framework that combines a shared voxel-map world model with a multi-agent Soft Actor-Critic (MASAC) controller. Multiple drones fuse 360 LiDAR observations into a common world-frame occupancy map, which is converted into a compact bird's-eye-view (BEV) representation and provided to each agent as an ego-aligned local crop. This integrate-in-world, act-in- ego design enables consistent multi-UAV spatial fusion whilst retaining decentralised co
arXiv 2d ago Research Agents & autonomy

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small task-relevant subset from a library of thousands of tools before the agent acts, has therefore become a critical component of LLM agent pipelines. However, existing retrievers either score each tool in isolation or assemble the tool set sequentially, so the joint utility of a candidate set is never evaluated as a whole. In this paper, we propose HYSET
arXiv cs.AI 2d ago Research Agents & autonomy

Token-maxing is an AI cost sink - how to use agents without busting your budget

Professionals are burning through tokens, but smart business leaders are finding ways to balance costs and value creation.
ZDNet AI 2d ago News Agents & autonomy

Microsoft dévoile Project Perception, sa propre armée d’agents IA pour la cybersécurité

Microsoft a présenté le 27 juillet 2026 une architecture de sécurité entièrement bâtie sur l'IA agentique, avec une préversion publique annoncée pour le 3 août 2026.
Numerama (FR) 2d ago News Agents & autonomy

Gemini Robotics 2 brings whole body intelligence to robots

Google DeepMind 2d ago Field notes Agents & autonomy

Tines seeks to tame ‘wild code’ AI sprawl

Workflow automation company Tines Ltd. today launched Tines 3B, an artificial intelligence-native platform for building, running and governing enterprise workflows, applications and agents. The product targets a problem Tines has taken to calling “Wild Code,” the sprawl of AI-generated software now being assembled inside enterprises by employees working outside traditional information technology processes. Much of […] The post Tines seeks to tame ‘wild code’ AI sprawl appeared first on SiliconAN
SiliconANGLE AI 2d ago News Jobs & economyAgents & autonomy

Snowflake debuts Cortex AI Gateway to govern and monitor enterprise AI agents

Snowflake Inc. today introduced Cortex AI Gateway, a centralized control layer that lets enterprises connect, govern and monitor artificial intelligence agents as they reach into models, tools, Model Context Protocol servers and internal systems. The company cast the gateway as the connective tissue for the “agentic enterprise,” the point at which autonomous software agents start […] The post Snowflake debuts Cortex AI Gateway to govern and monitor enterprise AI agents appeared first on SiliconA
SiliconANGLE AI 2d ago News Agents & autonomy

Exclusive: Dymium introduces single gateway to govern enterprise AI use

Secure artificial intelligence infrastructure startup Dymium Inc. today introduced GhostAI, a gateway designed to apply security and governance policies across the models, data, context and tools used in enterprise AI systems. The company said GhostAI sits between enterprise data and the AI models, agents and tools seeking to access it. The gateway inspects each interaction, […] The post Exclusive: Dymium introduces single gateway to govern enterprise AI use appeared first on SiliconANGLE .
SiliconANGLE AI 2d ago News RegulationAgents & autonomy

Diagrid Catalyst 2.0 adds durable execution to more than 10 agent frameworks

Agent infrastructure startup Diagrid Inc. today released Catalyst 2.0, an update to its managed workflow engine that adds automatic failure recovery and cryptographic verification to artificial intelligence agents built on frameworks such a LangGraph, Microsoft Agent Framework and Google’s Agent Development Kit. Developers do not have to rebuild anything to use it. Teams add a […] The post Diagrid Catalyst 2.0 adds durable execution to more than 10 agent frameworks appeared first on SiliconANGLE
SiliconANGLE AI 2d ago News Agents & autonomy

Diagrid gives failed AI agents a way to resume

AI agents can impress in a demo and still fumble in production. Diagrid’s Catalyst 2.0 aims to make them more The post Diagrid gives failed AI agents a way to resume appeared first on The New Stack .
The New Stack AI 2d ago News Agents & autonomy

AUTOBACS SEVEN Marks 10 Years of System Stability and Self-Funded Innovation with Rimini Street

Rimini Street, Inc. (Nasdaq: RMNI), the Software Support and Agentic AI ERP Company™ and the leading third-party support provider for Oracle, SAP and VMware software, today announced that AUTOBACS SEVEN Co., Ltd. celebrates its 10-year partnership with Rimini Street, marking a decade of stable core operations and reinvestment in innovation.
ITWeb (ZA) 2d ago News Agents & autonomy

Perplexity’s Personal Computer turns Windows PCs into AI agents

Perplexity has expanded its agentic Personal Computer tool to Windows, allowing computers running the world's most popular OS to be used as a locally run AI system. Like the Mac version that Perplexity launched in April, Personal Computer for Windows operates like a "general-purpose digital worker" that can access local files and apps to perform […]
The Verge 2d ago News Agents & autonomy

Lücke: Claude Cowork entkommt macOS-Sandbox

Die Nutzung von KI-Agenten direkt auf dem Rechner kann Gefahren mit sich bringen. Das zeigt eine soeben entdecktes Sicherheitsloch in Claude Cowork für den Mac.
Heise Online (DE) 2d ago News Agents & autonomy

Trulioo launches AI agent for beneficial ownership registry

Trulioo, a global risk intelligence platform, today announced the UBO Discovery Agent, the newest layer in Trulioo’s UBO Discovery capability inside its business risk and Know Your Business (KYB) verification workflow.
Finextra AI 2d ago News Agents & autonomy

(g+) Softwarentwicklung mit KI-Agenten: Die Slop-Maschine kontrollieren

Wie ich eine KI dazu bringe, guten Code zu schreiben - ganz praktisch. Ein Erfahrungsbericht von Felix Knorr ( KI , Softwareentwicklung )
Golem (DE) 2d ago News Agents & autonomy

How do we prevent AI agents from going rogue? It starts with a new kind of measurement | Bruce Schneier and Barath Raghavan

Like genies of folklore, AI agents take their instructions literally – to potentially disastrous effect. We must track their ability to do what we actually mean In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever was behind it captured internal security credentials and moved through systems over a weekend, running thousands of actions from a swarm of tempora
The Guardian 2d ago News Agents & autonomy

Meet the new robot dog patrolling LaGuardia Airport

A wheeled bot that monitors air quality is joining robotic vacuums and scrubbers in Terminal B.
Fast Company 2d ago News Agents & autonomy

Meet the new robot dog patrolling LaGuardia Airport

Inside the slick new Terminal B at New York’s LaGuardia Airport, a headless robotic “dog” is beginning to patrol the floors. As bemused airport travelers watch from close by, the four-wheeled device rolls around baggage claim, where it’s been deployed for a demonstration. Its job: sniffing. Well, sort of. The robot is armed with air quality detectors that the airport team says help monitor the terminal for pollutants.  This robot, along with two others, constitutes a fledgling automated fle
Fast Company Tech 2d ago News Jobs & economyAgents & autonomy

Meet the new robot dog patrolling LaGuardia Airport

A wheeled bot that monitors air quality is joining robotic vacuums and scrubbers in Terminal B.
Fast Company 2d ago News Agents & autonomy

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation mode
arXiv 2d ago Research HealthcareAgents & autonomy

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol). IMACS (Intelligent Multi-Agent Collaboration System) separates the three into orthogonal, independently swappable layers. Classic organizational theory (Belbin roles, Mintzberg coordination, RACI accountability) becomes executable, validated configurati
arXiv 2d ago Research Agents & autonomyTransparency

Crosby to Insure Its Agents for Legal Liability

NewMod law firm Crosby is to provide ‘professional liability insurance for [their] agents, so that they can do autonomous legal work’, in what is an ...
Artificial Lawyer 2d ago News RegulationAgents & autonomy

Agentic AI Autonomy Assessment: A Decision-Support Framework Towards Governed Supply Chain Systems

Supply chain decision-making is rapidly transforming with the rise of agentic AI - highly autonomous systems that can operate on complex, long-horizon tasks. Yet the adoption of agentic systems outpaces their governance: existing taxonomies of autonomy only offer discrete classifications, rely on subjective judgement, and cannot track autonomy across a system's life cycle, leaving enterprises unable to assess the risks of increasingly autonomous supply chain agents. This paper proposes the Agent
arXiv cs.HC 2d ago Research RegulationAgents & autonomy

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify not only what outcome to achieve, but also which steps, branches, and tool interactions are permitted. When these instructions are supplied as prompt context, however, the model retains control over both procedure selection and step execution. As interactions accumulate, an agent can skip required steps, take unsupported branches, or execute a vali
arXiv 2d ago Research Agents & autonomy

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon. We present HANDBOOK.md, a benchmark of 65 agentic
arXiv 2d ago Research RegulationAgents & autonomy

Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline authority: conflicts deny, incomplete or uncertain cl
arXiv 2d ago Research RegulationAgents & autonomy

AWS Launches Amazon GuardDuty Investigation Agent to Automate Threat Triage

AWS released a public preview of the GuardDuty investigation agent, which correlates findings, 90-day activity logs, and resource topologies into structured reports with risk ratings, confidence scores, and MITRE ATT&CK classification. It is reachable through the AWS MCP Server, so investigations can run from agentic tooling. Preview quotas cap usage at 10 investigations per account per day. By Steef-Jan Wiggers
InfoQ AI/ML 2d ago News Agents & autonomyFinance, VC & PE

ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design

This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes three contributions. First, it introduces a structured contract as the semantic-alignment and task-execution artifact that translates natural language requirements into explicit interfaces, constraints, validation checks, and rollback rules. Second, it incorporates hardware information into the feedback loop by feeding HLS, Vivado, PYNQ runtime, po
arXiv 3d ago Research Safety & alignmentAgents & autonomy

The Clinical Trial Pipeline Reveals the Next Wave of Artificial Intelligence in Healthcare: A Multidimensional Analysis of 8,532 Registered Studies

arXiv:2607.22607v1 Announce Type: new Abstract: The prospective clinical evaluation of artificial intelligence in medicine has expanded rapidly, but the global AI clinical trial landscape remains incompletely characterized. We systematically identified AI-related trials registered in ClinicalTrials.gov using a broad keyword search followed by an LLM-based classifier. Each trial was classified across seven dimensions: clinical function, data modality, specialty, AI integration and autonomy, workf
arXiv cs.CY 3d ago Research HealthcareAgents & autonomy

Accountable yet Anonymous AI Agents - Split-Knowledge Binding in National Agent-Identity Layer in China

arXiv:2607.23207v1 Announce Type: new Abstract: The emerging infrastructure for AI-agent identity has converged, in industry practice and research proposals alike, on a single resolution of the tension between accountability and privacy: make every agent identifiable. We document a national system in China -- built as national infrastructure and scheduled for public launch in Q3 2026 -- that occupies a different and underexplored point in the same design space: an agent is associated with a veri
arXiv cs.CY 3d ago Research PrivacyAgents & autonomy

Constitutional governance for societies of AI agents in the built environment: a research agenda

arXiv:2607.23336v1 Announce Type: new Abstract: The built environment is on the cusp of populating itself with autonomous artificial agents. AI systems that advise, control and coordinate are being deployed across retrofit, operation and mobility faster than their collective behaviour is studied. The dominant framing treats each agent as a tool operating on a passive building, governance reduced to single-agent safety, which is inadequate. A building, a street, or a city is more accurately model
arXiv cs.CY 3d ago Research RegulationAgents & autonomy

Private Again: AI Agents Restore Anonymity---Foreclosing Discrimination and Its Proof

arXiv:2607.23539v1 Announce Type: new Abstract: AI agents can transact online on behalf of a human principal---browsing, paying, receiving, and reviewing---without linking a transaction to a principal. That architecture starves algorithmic discrimination of its inputs---identity, purchase history, location history, behavioral traces, and demographic proxies---but also forecloses its proof. Disparate-treatment needs comparators; disparate-impact needs protected-class baselines; and Iqbal-era plea
arXiv cs.CY 3d ago Research Bias & fairnessAgents & autonomy

State-dependent error correlations shape voting thresholds in committees of AI agents

arXiv:2607.23931v1 Announce Type: new Abstract: The aggregation benefit of a committee of artificial intelligence (AI) agents comes from complementary information across members. Classical voting guarantees assume independent errors. Language-model errors often co-occur on the same cases. We combine Sah-Stiglitz screening with error dependence that can differ between good and bad cases. In a homogeneous exchangeable Gaussian-copula model, shared errors create a positive asymptotic error floor fo
arXiv cs.CY 3d ago Research Agents & autonomy

Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI

arXiv:2607.22953v1 Announce Type: cross Abstract: Modern AI systems bring societal risks such as mass surveillance, extreme concentrations of power, and loss of user autonomy---calling into question a model where third-parties collect and control massive amounts of user data. Users require a sovereign system to securely own, govern, and disclose their context while remaining compliant across regulated domains with strict provenance, interpretability, and policy adherence. Perspective-aware AI ap
arXiv cs.CY 3d ago Research RegulationSafety & alignment

Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels

arXiv:2607.23438v1 Announce Type: cross Abstract: As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they should be permitted to do in practice. This paper introduces a governance framework that explicitly separates Allowed Autonomy Levels (AAL), which define the degree of autonomy an AI agent is authorized to exercise given risk, oversight, and accountability considerations, from Autonomous Capabili
arXiv cs.CY 3d ago Research RegulationAgents & autonomy

Mapping the Reddit Bot Ecosystem: Taxonomy and Evolution

arXiv:2607.23941v1 Announce Type: cross Abstract: Automated agents increasingly participate in online communities, yet their population structure and roles remain poorly understood. Using a dataset of 3,389 identified bots and their full activity histories, we construct a taxonomy of bot "species" on the news aggregation and social media platform Reddit based on temporal, community, linguistic, and semantic features. Clustering analysis reveals 18 distinct bot types spanning content-specialized,
arXiv cs.CY 3d ago Research Agents & autonomy

Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure

arXiv:2603.28371v2 Announce Type: replace Abstract: When an agent can articulate why something works, we typically take this as evidence of genuine understanding. This presupposes that effective action and correct explanation covary, and that coherent explanation reliably signals both. I argue that this assumption fails for contemporary Large Language Models (LLMs). I introduce what I call the Bidirectional Coherence Paradox: competence and grounding not only dissociate but invert across epistem
arXiv cs.CY 3d ago Research Agents & autonomy
← Newer Older →