Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
Scientific computing in the age of agentic AI
A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and beyond.
Who is scientific code for? Maintaining human-readable landmarks in agent-written code
Scientific research involving code has long rested on the assumption that at least one person understands why the code exists. As scientists adopt coding agents, this assumption is breaking down. Drawing on an ongoing contextual inquiry of scientific programmers working with agentic tools (four cases to date), a survey of over 800 scientific programmers, and my own analysis workflows, this position piece describes how scientists are inventing personal conventions, "landmarking strategies", for m
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annotation pipeline provides frame-level clipping labels. This setup allows us to evaluate recent VLMs on a controlled anomaly detection task without manual annotation. We benchmark six recent VLMs (Gemini
Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks
Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Vendor B is compromised, agents from Vendor A continue invoking it -- unaware of the trust degradation -- causing cascading service impact. We present AgentToolMO, a proposed 3GPP NRM information model for agent tool trust management. The model comprises: a formally def
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $π_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we loc
GPT-Red: Automated Red Teaming via Self-Play at Scale
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the sc
Gemini API Managed Agents: 3.6 Flash, hooks, and more
We’re announcing even more new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six und
Distributing Security Controls Through Harness Engineering
AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI across organizations. Existing security controls for coding agents are not systematically distributed to engineering teams, and vendor-native solutions introduce ecosystem dependencies that may not suit every deployment context. This paper investigates whether off-the-shelf security controls can be implemented on commercial AI coding agents and scaled to a distrib
Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks
This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focus on uncertainty quantification. Actuarial workflows represent a high-stakes decision-support setting where unreliable outputs may lead to incorrect risk assessment, unfair pricing, and regulatory non-compliance. To address uncertainty introduced by the probabilistic nature of LLMs and dependencies between agents, a multi-agent framework is propo
HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs
Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. However, existing trajectory-to-skill methods often produce flat collections of high-level textual skills that are stored and retrieved independently, leaving skill relations underutilized and maintaining a gap between high-level skills and executable actions. In this paper, we propose HiSkill, a hierarchical skill graph framework that organizes i
Corpay introduces agent card capability
Corpay, a global leader in corporate payments, today announced Agent Card, a new capability that enables secure virtual card creation for AI-driven commerce workflows.
Fenergo launches Fen-AI to bring governed AI to client lifecycle management
Fenergo, the leading provider of digital solutions for Know Your Customer (KYC), Anti-Money Laundering (AML) and client lifecycle management (CLM), today announced the launch of Fen-AI, its agentic AI orchestration platform for financial institutions.
Lowering the implementation barrier of neutral-atom quantum computing with agentic workflows
Quantum computers are moving from research laboratories to industrial machines accessible via the cloud and integrated into high-performance computing facilities. However, translating theoretical quantum protocols into hardware experiments remains a major bottleneck, requiring expertise across protocol design, compilation, simulation, and cloud execution. Here, we introduce an agentic workflow that automates this pipeline on neutral-atom quantum processors (here two Pasqal QPUs available on the
Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predicting and pre-executing an agent's next tool call if the prediction matches the agent's eventual tool call, but existing speculators are typically separate draft models or cached traces that are poorly aligned with the deployed agent's own behavior. We identify this speculator-agent gap and show that the target agent itself is a strong next-call s
Wissenschaftler testen Ernteroboter auf Obstplantage am Bodensee
Am Bodensee wird ein Roboter getestet, der für den Einsatz im Obstanbau entwickelt wurde. Er soll künftig bei der Apfelernte eingesetzt werden.
GSA inks agentic AI OneGov deal with CORAS
CORAS is a FedRAMP High-certified agentic-AI “decision maker” that has already been authorized for use at the Department of Defense. The post GSA inks agentic AI OneGov deal with CORAS appeared first on FedScoop .
Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Critic Controller
This paper presents a cooperative indoor UAV guidance framework that combines a shared voxel-map world model with a multi-agent Soft Actor-Critic (MASAC) controller. Multiple drones fuse 360 LiDAR observations into a common world-frame occupancy map, which is converted into a compact bird's-eye-view (BEV) representation and provided to each agent as an ego-aligned local crop. This integrate-in-world, act-in- ego design enables consistent multi-UAV spatial fusion whilst retaining decentralised co
Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small task-relevant subset from a library of thousands of tools before the agent acts, has therefore become a critical component of LLM agent pipelines. However, existing retrievers either score each tool in isolation or assemble the tool set sequentially, so the joint utility of a candidate set is never evaluated as a whole. In this paper, we propose HYSET
Token-maxing is an AI cost sink - how to use agents without busting your budget
Professionals are burning through tokens, but smart business leaders are finding ways to balance costs and value creation.
Microsoft dévoile Project Perception, sa propre armée d’agents IA pour la cybersécurité
Microsoft a présenté le 27 juillet 2026 une architecture de sécurité entièrement bâtie sur l'IA agentique, avec une préversion publique annoncée pour le 3 août 2026.
Diagrid gives failed AI agents a way to resume
AI agents can impress in a demo and still fumble in production. Diagrid’s Catalyst 2.0 aims to make them more The post Diagrid gives failed AI agents a way to resume appeared first on The New Stack .
AUTOBACS SEVEN Marks 10 Years of System Stability and Self-Funded Innovation with Rimini Street
Rimini Street, Inc. (Nasdaq: RMNI), the Software Support and Agentic AI ERP Company™ and the leading third-party support provider for Oracle, SAP and VMware software, today announced that AUTOBACS SEVEN Co., Ltd. celebrates its 10-year partnership with Rimini Street, marking a decade of stable core operations and reinvestment in innovation.
Perplexity’s Personal Computer turns Windows PCs into AI agents
Perplexity has expanded its agentic Personal Computer tool to Windows, allowing computers running the world's most popular OS to be used as a locally run AI system. Like the Mac version that Perplexity launched in April, Personal Computer for Windows operates like a "general-purpose digital worker" that can access local files and apps to perform […]
Lücke: Claude Cowork entkommt macOS-Sandbox
Die Nutzung von KI-Agenten direkt auf dem Rechner kann Gefahren mit sich bringen. Das zeigt eine soeben entdecktes Sicherheitsloch in Claude Cowork für den Mac.
Trulioo launches AI agent for beneficial ownership registry
Trulioo, a global risk intelligence platform, today announced the UBO Discovery Agent, the newest layer in Trulioo’s UBO Discovery capability inside its business risk and Know Your Business (KYB) verification workflow.
(g+) Softwarentwicklung mit KI-Agenten: Die Slop-Maschine kontrollieren
Wie ich eine KI dazu bringe, guten Code zu schreiben - ganz praktisch. Ein Erfahrungsbericht von Felix Knorr ( KI , Softwareentwicklung )
How do we prevent AI agents from going rogue? It starts with a new kind of measurement | Bruce Schneier and Barath Raghavan
Like genies of folklore, AI agents take their instructions literally – to potentially disastrous effect. We must track their ability to do what we actually mean In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever was behind it captured internal security credentials and moved through systems over a weekend, running thousands of actions from a swarm of tempora
Meet the new robot dog patrolling LaGuardia Airport
A wheeled bot that monitors air quality is joining robotic vacuums and scrubbers in Terminal B.
Meet the new robot dog patrolling LaGuardia Airport
Inside the slick new Terminal B at New York’s LaGuardia Airport, a headless robotic “dog” is beginning to patrol the floors. As bemused airport travelers watch from close by, the four-wheeled device rolls around baggage claim, where it’s been deployed for a demonstration. Its job: sniffing. Well, sort of. The robot is armed with air quality detectors that the airport team says help monitor the terminal for pollutants. This robot, along with two others, constitutes a fledgling automated fle
Meet the new robot dog patrolling LaGuardia Airport
A wheeled bot that monitors air quality is joining robotic vacuums and scrubbers in Terminal B.
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation mode
Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm
Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol). IMACS (Intelligent Multi-Agent Collaboration System) separates the three into orthogonal, independently swappable layers. Classic organizational theory (Belbin roles, Mintzberg coordination, RACI accountability) becomes executable, validated configurati
Crosby to Insure Its Agents for Legal Liability
NewMod law firm Crosby is to provide ‘professional liability insurance for [their] agents, so that they can do autonomous legal work’, in what is an ...
Agentic AI Autonomy Assessment: A Decision-Support Framework Towards Governed Supply Chain Systems
Supply chain decision-making is rapidly transforming with the rise of agentic AI - highly autonomous systems that can operate on complex, long-horizon tasks. Yet the adoption of agentic systems outpaces their governance: existing taxonomies of autonomy only offer discrete classifications, rely on subjective judgement, and cannot track autonomy across a system's life cycle, leaving enterprises unable to assess the risks of increasingly autonomous supply chain agents. This paper proposes the Agent
COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution
Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify not only what outcome to achieve, but also which steps, branches, and tool interactions are permitted. When these instructions are supplied as prompt context, however, the model retains control over both procedure selection and step execution. As interactions accumulate, an agent can skip required steps, take unsupported branches, or execute a vali
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon. We present HANDBOOK.md, a benchmark of 65 agentic
Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline authority: conflicts deny, incomplete or uncertain cl
AWS Launches Amazon GuardDuty Investigation Agent to Automate Threat Triage
AWS released a public preview of the GuardDuty investigation agent, which correlates findings, 90-day activity logs, and resource topologies into structured reports with risk ratings, confidence scores, and MITRE ATT&CK classification. It is reachable through the AWS MCP Server, so investigations can run from agentic tooling. Preview quotas cap usage at 10 investigations per account per day. By Steef-Jan Wiggers