02:30 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

Scientific computing in the age of agentic AI

A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and beyond.
OpenAI 2d ago Field notes Agents & autonomyBiotech

Who is scientific code for? Maintaining human-readable landmarks in agent-written code

Scientific research involving code has long rested on the assumption that at least one person understands why the code exists. As scientists adopt coding agents, this assumption is breaking down. Drawing on an ongoing contextual inquiry of scientific programmers working with agentic tools (four cases to date), a survey of over 800 scientific programmers, and my own analysis workflows, this position piece describes how scientists are inventing personal conventions, "landmarking strategies", for m
arXiv cs.HC 2d ago Research Agents & autonomy

Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA

In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annotation pipeline provides frame-level clipping labels. This setup allows us to evaluate recent VLMs on a controlled anomaly detection task without manual annotation. We benchmark six recent VLMs (Gemini
arXiv cs.AI 2d ago Research Agents & autonomy

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Vendor B is compromised, agents from Vendor A continue invoking it -- unaware of the trust degradation -- causing cascading service impact. We present AgentToolMO, a proposed 3GPP NRM information model for agent tool trust management. The model comprises: a formally def
arXiv cs.AI 2d ago Research Agents & autonomy

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $π_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we loc
arXiv 2d ago Research Safety & alignmentAgents & autonomy

GPT-Red: Automated Red Teaming via Self-Play at Scale

We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender
arXiv red teaming query 2d ago Research Safety & alignmentAgents & autonomy

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the sc
arXiv cs.AI 2d ago Research Agents & autonomyEnvironment

Gemini API Managed Agents: 3.6 Flash, hooks, and more

We’re announcing even more new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.
Google AI Blog 2d ago Field notes Agents & autonomy

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six und
arXiv 2d ago Research Agents & autonomyEnvironment

Distributing Security Controls Through Harness Engineering

AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI across organizations. Existing security controls for coding agents are not systematically distributed to engineering teams, and vendor-native solutions introduce ecosystem dependencies that may not suit every deployment context. This paper investigates whether off-the-shelf security controls can be implemented on commercial AI coding agents and scaled to a distrib
arXiv cs.AI 2d ago Research Agents & autonomyFinance, VC & PE

Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focus on uncertainty quantification. Actuarial workflows represent a high-stakes decision-support setting where unreliable outputs may lead to incorrect risk assessment, unfair pricing, and regulatory non-compliance. To address uncertainty introduced by the probabilistic nature of LLMs and dependencies between agents, a multi-agent framework is propo
arXiv cs.AI 2d ago Research RegulationAgents & autonomy

HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs

Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. However, existing trajectory-to-skill methods often produce flat collections of high-level textual skills that are stored and retrieved independently, leaving skill relations underutilized and maintaining a gap between high-level skills and executable actions. In this paper, we propose HiSkill, a hierarchical skill graph framework that organizes i
arXiv cs.AI 2d ago Research Agents & autonomy

Corpay introduces agent card capability

Corpay, a global leader in corporate payments, today announced Agent Card, a new capability that enables secure virtual card creation for AI-driven commerce workflows.
Finextra AI 2d ago News Agents & autonomy

Fenergo launches Fen-AI to bring governed AI to client lifecycle management

Fenergo, the leading provider of digital solutions for Know Your Customer (KYC), Anti-Money Laundering (AML) and client lifecycle management (CLM), today announced the launch of Fen-AI, its agentic AI orchestration platform for financial institutions.
Finextra AI 2d ago News Agents & autonomy

Lowering the implementation barrier of neutral-atom quantum computing with agentic workflows

Quantum computers are moving from research laboratories to industrial machines accessible via the cloud and integrated into high-performance computing facilities. However, translating theoretical quantum protocols into hardware experiments remains a major bottleneck, requiring expertise across protocol design, compilation, simulation, and cloud execution. Here, we introduce an agentic workflow that automates this pipeline on neutral-atom quantum processors (here two Pasqal QPUs available on the
arXiv cs.AI 2d ago Research Agents & autonomy

Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL

Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predicting and pre-executing an agent's next tool call if the prediction matches the agent's eventual tool call, but existing speculators are typically separate draft models or cached traces that are poorly aligned with the deployed agent's own behavior. We identify this speculator-agent gap and show that the target agent itself is a strong next-call s
arXiv 2d ago Research Agents & autonomy

Wissenschaftler testen Ernteroboter auf Obstplantage am Bodensee

Am Bodensee wird ein Roboter getestet, der für den Einsatz im Obstanbau entwickelt wurde. Er soll künftig bei der Apfelernte eingesetzt werden.
Heise Online (DE) 2d ago News Agents & autonomy

GSA inks agentic AI OneGov deal with CORAS

CORAS is a FedRAMP High-certified agentic-AI “decision maker” that has already been authorized for use at the Department of Defense. The post GSA inks agentic AI OneGov deal with CORAS appeared first on FedScoop .
FedScoop 2d ago News Military & securityAgents & autonomy

Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Critic Controller

This paper presents a cooperative indoor UAV guidance framework that combines a shared voxel-map world model with a multi-agent Soft Actor-Critic (MASAC) controller. Multiple drones fuse 360 LiDAR observations into a common world-frame occupancy map, which is converted into a compact bird's-eye-view (BEV) representation and provided to each agent as an ego-aligned local crop. This integrate-in-world, act-in- ego design enables consistent multi-UAV spatial fusion whilst retaining decentralised co
arXiv 2d ago Research Agents & autonomy

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small task-relevant subset from a library of thousands of tools before the agent acts, has therefore become a critical component of LLM agent pipelines. However, existing retrievers either score each tool in isolation or assemble the tool set sequentially, so the joint utility of a candidate set is never evaluated as a whole. In this paper, we propose HYSET
arXiv cs.AI 2d ago Research Agents & autonomy

Token-maxing is an AI cost sink - how to use agents without busting your budget

Professionals are burning through tokens, but smart business leaders are finding ways to balance costs and value creation.
ZDNet AI 2d ago News Agents & autonomy

Microsoft dévoile Project Perception, sa propre armée d’agents IA pour la cybersécurité

Microsoft a présenté le 27 juillet 2026 une architecture de sécurité entièrement bâtie sur l'IA agentique, avec une préversion publique annoncée pour le 3 août 2026.
Numerama (FR) 2d ago News Agents & autonomy

Diagrid gives failed AI agents a way to resume

AI agents can impress in a demo and still fumble in production. Diagrid’s Catalyst 2.0 aims to make them more The post Diagrid gives failed AI agents a way to resume appeared first on The New Stack .
The New Stack AI 2d ago News Agents & autonomy

AUTOBACS SEVEN Marks 10 Years of System Stability and Self-Funded Innovation with Rimini Street

Rimini Street, Inc. (Nasdaq: RMNI), the Software Support and Agentic AI ERP Company™ and the leading third-party support provider for Oracle, SAP and VMware software, today announced that AUTOBACS SEVEN Co., Ltd. celebrates its 10-year partnership with Rimini Street, marking a decade of stable core operations and reinvestment in innovation.
ITWeb (ZA) 2d ago News Agents & autonomy

Perplexity’s Personal Computer turns Windows PCs into AI agents

Perplexity has expanded its agentic Personal Computer tool to Windows, allowing computers running the world's most popular OS to be used as a locally run AI system. Like the Mac version that Perplexity launched in April, Personal Computer for Windows operates like a "general-purpose digital worker" that can access local files and apps to perform […]
The Verge 2d ago News Agents & autonomy

Lücke: Claude Cowork entkommt macOS-Sandbox

Die Nutzung von KI-Agenten direkt auf dem Rechner kann Gefahren mit sich bringen. Das zeigt eine soeben entdecktes Sicherheitsloch in Claude Cowork für den Mac.
Heise Online (DE) 2d ago News Agents & autonomy

Trulioo launches AI agent for beneficial ownership registry

Trulioo, a global risk intelligence platform, today announced the UBO Discovery Agent, the newest layer in Trulioo’s UBO Discovery capability inside its business risk and Know Your Business (KYB) verification workflow.
Finextra AI 2d ago News Agents & autonomy

(g+) Softwarentwicklung mit KI-Agenten: Die Slop-Maschine kontrollieren

Wie ich eine KI dazu bringe, guten Code zu schreiben - ganz praktisch. Ein Erfahrungsbericht von Felix Knorr ( KI , Softwareentwicklung )
Golem (DE) 2d ago News Agents & autonomy

How do we prevent AI agents from going rogue? It starts with a new kind of measurement | Bruce Schneier and Barath Raghavan

Like genies of folklore, AI agents take their instructions literally – to potentially disastrous effect. We must track their ability to do what we actually mean In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever was behind it captured internal security credentials and moved through systems over a weekend, running thousands of actions from a swarm of tempora
The Guardian 2d ago News Agents & autonomy

Meet the new robot dog patrolling LaGuardia Airport

A wheeled bot that monitors air quality is joining robotic vacuums and scrubbers in Terminal B.
Fast Company 2d ago News Agents & autonomy

Meet the new robot dog patrolling LaGuardia Airport

Inside the slick new Terminal B at New York’s LaGuardia Airport, a headless robotic “dog” is beginning to patrol the floors. As bemused airport travelers watch from close by, the four-wheeled device rolls around baggage claim, where it’s been deployed for a demonstration. Its job: sniffing. Well, sort of. The robot is armed with air quality detectors that the airport team says help monitor the terminal for pollutants.  This robot, along with two others, constitutes a fledgling automated fle
Fast Company Tech 2d ago News Jobs & economyAgents & autonomy

Meet the new robot dog patrolling LaGuardia Airport

A wheeled bot that monitors air quality is joining robotic vacuums and scrubbers in Terminal B.
Fast Company 2d ago News Agents & autonomy

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation mode
arXiv 2d ago Research HealthcareAgents & autonomy

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol). IMACS (Intelligent Multi-Agent Collaboration System) separates the three into orthogonal, independently swappable layers. Classic organizational theory (Belbin roles, Mintzberg coordination, RACI accountability) becomes executable, validated configurati
arXiv 2d ago Research Agents & autonomyTransparency

Crosby to Insure Its Agents for Legal Liability

NewMod law firm Crosby is to provide ‘professional liability insurance for [their] agents, so that they can do autonomous legal work’, in what is an ...
Artificial Lawyer 2d ago News RegulationAgents & autonomy

Agentic AI Autonomy Assessment: A Decision-Support Framework Towards Governed Supply Chain Systems

Supply chain decision-making is rapidly transforming with the rise of agentic AI - highly autonomous systems that can operate on complex, long-horizon tasks. Yet the adoption of agentic systems outpaces their governance: existing taxonomies of autonomy only offer discrete classifications, rely on subjective judgement, and cannot track autonomy across a system's life cycle, leaving enterprises unable to assess the risks of increasingly autonomous supply chain agents. This paper proposes the Agent
arXiv cs.HC 2d ago Research RegulationAgents & autonomy

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify not only what outcome to achieve, but also which steps, branches, and tool interactions are permitted. When these instructions are supplied as prompt context, however, the model retains control over both procedure selection and step execution. As interactions accumulate, an agent can skip required steps, take unsupported branches, or execute a vali
arXiv 2d ago Research Agents & autonomy

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon. We present HANDBOOK.md, a benchmark of 65 agentic
arXiv 2d ago Research RegulationAgents & autonomy

Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline authority: conflicts deny, incomplete or uncertain cl
arXiv 2d ago Research RegulationAgents & autonomy

AWS Launches Amazon GuardDuty Investigation Agent to Automate Threat Triage

AWS released a public preview of the GuardDuty investigation agent, which correlates findings, 90-day activity logs, and resource topologies into structured reports with risk ratings, confidence scores, and MITRE ATT&CK classification. It is reachable through the AWS MCP Server, so investigations can run from agentic tooling. Preview quotas cap usage at 10 investigations per account per day. By Steef-Jan Wiggers
InfoQ AI/ML 2d ago News Agents & autonomyFinance, VC & PE
← Newer Older →