21:25 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can become trapped in repetitive loops, wasting search budgets and ultimately compromising the quality and completeness of the final output. We introduce SearchOS, a system-level multi-age
arXiv cs.AI 15d ago Research Agents & autonomy

GoDaddy opened its registrar to AI agents. Then it had to build guardrails.

On Wednesday, GoDaddy launched its new developer platform, giving developers a way to manage domains without leaving their development environment. The post GoDaddy opened its registrar to AI agents. Then it had to build guardrails. appeared first on The New Stack .
The New Stack AI 15d ago News Agents & autonomyEnvironment

AutoSynthesis: An agentic system for automated meta-analysis

Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quantitative evidence synthesis remains largely manual and difficult to scale. Here, we introduce AutoSynthesis, an end-to-end multi-agent system for automated meta-analysis. Given a research question in natural language, AutoSynthesis formulates a search strategy, retrieves scientific literature, screens candidate studies, assesses full-text eligibility, extracts
arXiv 15d ago Research RegulationChildren & education

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-state direction analysis and random-split null tests, we show that content danger (CD) and physical danger (PD) form separable signals in LLM representations across Qwen2.5-3B/7B/14B
arXiv cs.AI 15d ago Research Agents & autonomy

The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix

Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the default context source, and provider-native retrieval has quietly overtaken the dedicated vector databases that define the category — yet a majority of enterprises have already watched their agents produce confident, wrong answers traced to missing or inconsistent context. A governed semantic layer is emerging as the fi
VentureBeat 15d ago News Agents & autonomy

BadWAM: When World-Action Models Dream Right but Act Wrong

World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluatin
arXiv cs.LG 15d ago Research Safety & alignmentAgents & autonomy

Plover: Steering GUI Agents through Plan-Centric Interaction

Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users' ability to inspect, supervise, or correct system behavior. We present Plover, a pla
arXiv cs.AI 15d ago Research Jobs & economyAgents & autonomy

The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway

Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to prod
VentureBeat 15d ago News Safety & alignmentAgents & autonomy

How to Make an Invisible Drone

There are many words that I would never, ever use to describe a drone. Stealthy. Subtle. Whatever the opposite of obnoxious is. Much of this is because of the giant angry bee sound that drones tend to make, but it’s also the way that they look in flight: With uncannily linear movements and an even less canny ability to hover perfectly still, they tend to draw the eye as affronts to nature. In a paper presented this week at RSS 2026 in Sydney, roboticists from Northwestern University, Evanston, I
IEEE Spectrum 15d ago News Agents & autonomy

Scaling Behavior Foundation Model for Humanoid Robots

Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to f
arXiv cs.AI 15d ago Research Agents & autonomyEnvironment

NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval

Hugging Face Blog 15d ago Field notes Agents & autonomy

Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents

AI coding agents set up projects by reading documentation and installing the dependencies it lists, without verifying their names, sources, or known vulnerabilities. By editing only a README, requirements file, or Makefile, an attacker can redirect the agent to an untrusted registry, a known-vulnerable version, or a wrong-but-plausible name: documentation becomes a vector for code execution. We present the first systematic evaluation of package-install-time supply-chain attacks delivered through
arXiv cs.HC 15d ago Research Military & securityAgents & autonomy

Concept-Guided Spatial Regularization for World Models in Atari Pong

World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, while the world models themselves are rarely studied in isolation. We examine five representative visual world-model agents in Atari Pong: DreamerV3, DIAMOND, TWISTER, Simulus, and STORM. After reproducing their training pipelines and matching the reported agent performance, we freeze the learned world models and evaluate them with a closed-loop rollout diagnostic: a policy trained separately f
arXiv cs.AI 15d ago Research RegulationHealthcare

Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling

Dynamic object exchange between humans and robots remains a challenging problem due to uncertainty in perception, timing, and contact-rich interaction. Human-robot juggling represents a particularly demanding instance of this problem, requiring precise real-time coordination, predictive motion planning with feedback control, and robustness to variability in human motion. Enabling such skills is of interest for advancing physical human-robot interaction and shared autonomy. We present a real-time
arXiv cs.HC 15d ago Research Agents & autonomy

Yes, you can now order DoorDash from the command line

DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place orders from the terminal, marking another step toward software designed for AI agents instead of just humans.
TechCrunch 15d ago News Agents & autonomy

Litera Relaunches Its Brand Around A Unified Vision: One AI Agent Spanning The Practice and Business Of Law

The relaunch includes a new brand identity, the tagline ‘Raise The Bar,’ an advertising campaign, and a redesigned website. The post Litera Relaunches Its Brand Around A Unified Vision: One AI Agent Spanning The Practice and Business Of Law appeared first on Above the Law .
Above the Law (legal tech) 15d ago News RegulationAgents & autonomy

heise+ | KI in der Musik: Wenn Agenten selbst komponieren

Musikwissenschaftler Matthias Röder stellte einst eine Beethoven-Sinfonie mittels KI fertig. Im Interview spricht er darüber, was er für die Zukunft erwartet.
Heise Online (DE) 15d ago News Agents & autonomy

Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents

The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neutrality and helpfulness biases instilled by Reinforcement Learning from Human Feedback (RLHF) prevent them from sustaining steadfast partisan behaviour. We present a multi-agent framework that reconciles factual grounding with ideological alignment by combin
arXiv 15d ago Research RegulationSafety & alignment

Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents

The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neutrality and helpfulness biases instilled by Reinforcement Learning from Human Feedback (RLHF) prevent them from sustaining steadfast partisan behaviour. We present a multi-agent framework that reconciles factual grounding with ideological alignment by combin
arXiv 15d ago Research RegulationSafety & alignment

OpenAI’s GPT-Red automates prompt injection testing to harden AI agents

Now that AI agents are performing real tasks rather than just generating text, the old methods of manual security testing The post OpenAI’s GPT-Red automates prompt injection testing to harden AI agents appeared first on The New Stack .
The New Stack AI 15d ago News Agents & autonomy

BrainPilot: Automating Brain Discovery with Agentic Research

Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveying prior work to executing analyses and interpreting results in light of domain knowledge. AI agents promise to accelerate this process, but current agents lack domain expertise in brain science, may fabricate claims, drift during multi-step reasoning, and offer few defined point
arXiv cs.AI 15d ago Research Agents & autonomy

The state of agentic storefronts: How AI agents guide shoppers on frictionless, full-funnel journeys

This State of the Industry report, sponsored by Swap, explores how brands, retailers and agencies are adopting agentic commerce, including agentic storefronts, to enhance shopping experiences. Artificial intelligence is the new entry point in e-commerce, as more shoppers rely on AI-powered summaries and assistants to discover products. Beyond discovery and product recommendations, AI agents are […]
Digiday (AI/media) 15d ago News Agents & autonomy

Codex Micro: OpenAIs erste Hardware ist ein Tastenpad

Der Codex Micro soll das Steuern von KI-Agenten mittels Tasten erleichtern. Daf�r arbeitet OpenAI mit einem Tastaturhersteller zusammen. ( Peripherieger�te , Eingabeger�t )
Golem (DE) 15d ago News Agents & autonomy

OpenAI wants developers to stop typing commands and start using a joystick to control their AI agents

OpenAI and keyboard manufacturer Work Louder have unveiled the Codex Micro, a compact hardware controller designed for working with AI agents. The article OpenAI wants developers to stop typing commands and start using a joystick to control their AI agents appeared first on The Decoder .
The Decoder 15d ago News Agents & autonomy

ANet Patu-1: The Value of Connection in the Agent Network

The Internet taught us that the value of a network depends on \emph{how} its nodes connect: broadcast stars scale as $V\!\propto\!N$ (Sarnoff), fully-connected meshes as $N^2$ (Metcalfe), and group-forming networks as $2^{N}$ (Reed). We ask the analogous question for networks of AI agents. We model the net value of connection as a function of coordination-group size, derive from it the properties an optimal collaboration protocol must have, and introduce ANet Patu-1 -- a self-organizing consensu
arXiv 15d ago Research Agents & autonomy

Thredd joins Visa's Agentic Ready programme

Thredd, the AI-first issuer processing platform, today announced it has joined the Visa Agentic Ready programme, enabling issuers across Europe to participate in agent-initiated payments without rebuilding their payments infrastructure.
Finextra AI 15d ago News Agents & autonomy

LQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics Research

Lattice quantum chromodynamics (LQCD) provides a first-principles framework for computing hadronic observables, but its practical use remains limited by the substantial expertise required to turn research motivation into reliable computing workflows. Here we present \textsc{LQCDMaster}, a tool-augmented, skill-guided and domain-specialized scientific computing agent that converts natural-language LQCD research tasks into executable PyQUDA computing workflows, including measurement scripts, job-s
arXiv cs.AI 15d ago Research Jobs & economyAgents & autonomy

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction. However, existing agent benchmarks often focus on limited scenarios, tool ecosystems, or interaction formats, making it difficult to systematically characterize model capabilities across heterogeneous application settings. We introduce OmniaBench, a benchmark for evaluating general agents across
arXiv cs.AI 15d ago Research Agents & autonomy

🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences

Lila is betting that science, not the internet, is the last untapped source of training data. We went to find out what that actually looks like in a room full of robots.
Latent Space 15d ago Field notes Agents & autonomyEnvironment

1Password's new Agentic Mode lets Claude log into your accounts without seeing your credentials

1Password wants to solve one of AI's biggest practical problems: secure logins. Its Claude integration can enter passwords and MFA codes without exposing credentials to Anthropic or the model.
ZDNet AI 15d ago News Agents & autonomy

You can now grant Claude access to your 1Password credentials

1Password lets you use Claude agents for personal chores without exposing your credentials.
Engadget AI 15d ago News Agents & autonomy

“There are no laws, only suggestions”: What AI agents do with your instructions

Nine days into a twelve-day experiment with an AI coding agent, SaaStr founder Jason Lemkin had one instruction on file: The post “There are no laws, only suggestions”: What AI agents do with your instructions appeared first on The New Stack .
The New Stack AI 15d ago News Agents & autonomy

Exclusive: Marc Lore says Wonder is gearing up for an IPO after raising $650 million at a $9 billion valuation

Wonder, which owns Grubhub and Blue Apron, is betting robotics, AI, and a nationwide expansion can reshape restaurant delivery before going public.
Fortune AI 15d ago News Agents & autonomyFinance, VC & PE

We Can’t Monitor AI Agents at Scale. Here’s What It Will Take.

Tech Policy Press 15d ago News Agents & autonomy

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, code-check records, and a final report. Evaluations centered on question answering or script generation rarely verify this complete evidence chain and may therefore reward fluent outputs even when the underlying engineering workflow is incomplete, internally inconsistent, or non-executabl
arXiv 15d ago Research Agents & autonomy

Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control

Autonomous coding agents increasingly execute multi-step software work, but lifecycle states such as reviewed, tested, DONE, and ready-to-merge remain claims unless supported by current evidence. We present Proof-or-Stop Lifecycle Control, a method that permits lifecycle transitions only when fresh, tracked-source-state-bound, mechanically verifiable evidence satisfies the relevant gate. The method treats agent outputs as claims rather than lifecycle state, and uses proof operationally to mean g
arXiv cs.AI 15d ago Research Agents & autonomy

KI-Assistent Bob 2.0 modernisiert IBM Z und Power-Systeme

IBM Bob 2.0 soll auch Jahrzehnte alten Code verstehen und modernisieren können – mit KI-Agenten und Premium-Paketen für Z, i und Java.
Heise Online (DE) 15d ago News Agents & autonomy

AI Agents with Cloud Credentials Are Outrunning Billing Guardrails Built for Human-Speed Mistakes

A three-person agency received a $14,000 AWS bill in one day after attackers extracted static access keys and burned Claude invocations on Bedrock. Combined with May's DN42 incident, where an autonomous agent provisioned $6,531 of oversized infrastructure in 24 hours, practitioners warn that cloud billing lags roughly a day behind agent-speed spend. By Steef-Jan Wiggers
InfoQ AI/ML 15d ago News RegulationAgents & autonomy

Circus Commences Operations with Ukrainian Ground Forces

Circus SE (WKN: A2YN35 / ISIN: DE000A2YN355 / XETRA: CA1), today announces the commencement of live operations of its robotic-based troop supply technology with the 3rd Army Corps of the Ukrainian Ground Forces in the Kyiv area – marking the first ever use of autonomous meal supply systems within an active conflict environment.
ITWeb (ZA) 15d ago News Agents & autonomyEnvironment

Knife Capital backs local AI agent business

Cape Town-headquartered Cue secures an R82 million funding round, co-led by the venture capital firm and FAM Investments.
ITWeb (ZA) 15d ago News Agents & autonomyFinance, VC & PE
← Newer Older →