Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can become trapped in repetitive loops, wasting search budgets and ultimately compromising the quality and completeness of the final output. We introduce SearchOS, a system-level multi-age
GoDaddy opened its registrar to AI agents. Then it had to build guardrails.
On Wednesday, GoDaddy launched its new developer platform, giving developers a way to manage domains without leaving their development environment. The post GoDaddy opened its registrar to AI agents. Then it had to build guardrails. appeared first on The New Stack .
AutoSynthesis: An agentic system for automated meta-analysis
Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quantitative evidence synthesis remains largely manual and difficult to scale. Here, we introduce AutoSynthesis, an end-to-end multi-agent system for automated meta-analysis. Given a research question in natural language, AutoSynthesis formulates a search strategy, retrieves scientific literature, screens candidate studies, assesses full-text eligibility, extracts
When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-state direction analysis and random-split null tests, we show that content danger (CD) and physical danger (PD) form separable signals in LLM representations across Qwen2.5-3B/7B/14B
The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix
Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the default context source, and provider-native retrieval has quietly overtaken the dedicated vector databases that define the category — yet a majority of enterprises have already watched their agents produce confident, wrong answers traced to missing or inconsistent context. A governed semantic layer is emerging as the fi
BadWAM: When World-Action Models Dream Right but Act Wrong
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluatin
Plover: Steering GUI Agents through Plan-Centric Interaction
Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users' ability to inspect, supervise, or correct system behavior. We present Plover, a pla
The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to prod
How to Make an Invisible Drone
There are many words that I would never, ever use to describe a drone. Stealthy. Subtle. Whatever the opposite of obnoxious is. Much of this is because of the giant angry bee sound that drones tend to make, but it’s also the way that they look in flight: With uncannily linear movements and an even less canny ability to hover perfectly still, they tend to draw the eye as affronts to nature. In a paper presented this week at RSS 2026 in Sydney, roboticists from Northwestern University, Evanston, I
Scaling Behavior Foundation Model for Humanoid Robots
Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to f
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval
Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents
AI coding agents set up projects by reading documentation and installing the dependencies it lists, without verifying their names, sources, or known vulnerabilities. By editing only a README, requirements file, or Makefile, an attacker can redirect the agent to an untrusted registry, a known-vulnerable version, or a wrong-but-plausible name: documentation becomes a vector for code execution. We present the first systematic evaluation of package-install-time supply-chain attacks delivered through
Concept-Guided Spatial Regularization for World Models in Atari Pong
World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, while the world models themselves are rarely studied in isolation. We examine five representative visual world-model agents in Atari Pong: DreamerV3, DIAMOND, TWISTER, Simulus, and STORM. After reproducing their training pipelines and matching the reported agent performance, we freeze the learned world models and evaluate them with a closed-loop rollout diagnostic: a policy trained separately f
Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling
Dynamic object exchange between humans and robots remains a challenging problem due to uncertainty in perception, timing, and contact-rich interaction. Human-robot juggling represents a particularly demanding instance of this problem, requiring precise real-time coordination, predictive motion planning with feedback control, and robustness to variability in human motion. Enabling such skills is of interest for advancing physical human-robot interaction and shared autonomy. We present a real-time
Yes, you can now order DoorDash from the command line
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place orders from the terminal, marking another step toward software designed for AI agents instead of just humans.
Litera Relaunches Its Brand Around A Unified Vision: One AI Agent Spanning The Practice and Business Of Law
The relaunch includes a new brand identity, the tagline ‘Raise The Bar,’ an advertising campaign, and a redesigned website. The post Litera Relaunches Its Brand Around A Unified Vision: One AI Agent Spanning The Practice and Business Of Law appeared first on Above the Law .
heise+ | KI in der Musik: Wenn Agenten selbst komponieren
Musikwissenschaftler Matthias Röder stellte einst eine Beethoven-Sinfonie mittels KI fertig. Im Interview spricht er darüber, was er für die Zukunft erwartet.
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neutrality and helpfulness biases instilled by Reinforcement Learning from Human Feedback (RLHF) prevent them from sustaining steadfast partisan behaviour. We present a multi-agent framework that reconciles factual grounding with ideological alignment by combin
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neutrality and helpfulness biases instilled by Reinforcement Learning from Human Feedback (RLHF) prevent them from sustaining steadfast partisan behaviour. We present a multi-agent framework that reconciles factual grounding with ideological alignment by combin
OpenAI’s GPT-Red automates prompt injection testing to harden AI agents
Now that AI agents are performing real tasks rather than just generating text, the old methods of manual security testing The post OpenAI’s GPT-Red automates prompt injection testing to harden AI agents appeared first on The New Stack .
BrainPilot: Automating Brain Discovery with Agentic Research
Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveying prior work to executing analyses and interpreting results in light of domain knowledge. AI agents promise to accelerate this process, but current agents lack domain expertise in brain science, may fabricate claims, drift during multi-step reasoning, and offer few defined point
The state of agentic storefronts: How AI agents guide shoppers on frictionless, full-funnel journeys
This State of the Industry report, sponsored by Swap, explores how brands, retailers and agencies are adopting agentic commerce, including agentic storefronts, to enhance shopping experiences. Artificial intelligence is the new entry point in e-commerce, as more shoppers rely on AI-powered summaries and assistants to discover products. Beyond discovery and product recommendations, AI agents are […]
Codex Micro: OpenAIs erste Hardware ist ein Tastenpad
Der Codex Micro soll das Steuern von KI-Agenten mittels Tasten erleichtern. Daf�r arbeitet OpenAI mit einem Tastaturhersteller zusammen. ( Peripherieger�te , Eingabeger�t )
OpenAI wants developers to stop typing commands and start using a joystick to control their AI agents
OpenAI and keyboard manufacturer Work Louder have unveiled the Codex Micro, a compact hardware controller designed for working with AI agents. The article OpenAI wants developers to stop typing commands and start using a joystick to control their AI agents appeared first on The Decoder .
ANet Patu-1: The Value of Connection in the Agent Network
The Internet taught us that the value of a network depends on \emph{how} its nodes connect: broadcast stars scale as $V\!\propto\!N$ (Sarnoff), fully-connected meshes as $N^2$ (Metcalfe), and group-forming networks as $2^{N}$ (Reed). We ask the analogous question for networks of AI agents. We model the net value of connection as a function of coordination-group size, derive from it the properties an optimal collaboration protocol must have, and introduce ANet Patu-1 -- a self-organizing consensu
Thredd joins Visa's Agentic Ready programme
Thredd, the AI-first issuer processing platform, today announced it has joined the Visa Agentic Ready programme, enabling issuers across Europe to participate in agent-initiated payments without rebuilding their payments infrastructure.
LQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics Research
Lattice quantum chromodynamics (LQCD) provides a first-principles framework for computing hadronic observables, but its practical use remains limited by the substantial expertise required to turn research motivation into reliable computing workflows. Here we present \textsc{LQCDMaster}, a tool-augmented, skill-guided and domain-specialized scientific computing agent that converts natural-language LQCD research tasks into executable PyQUDA computing workflows, including measurement scripts, job-s
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction. However, existing agent benchmarks often focus on limited scenarios, tool ecosystems, or interaction formats, making it difficult to systematically characterize model capabilities across heterogeneous application settings. We introduce OmniaBench, a benchmark for evaluating general agents across
🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
Lila is betting that science, not the internet, is the last untapped source of training data. We went to find out what that actually looks like in a room full of robots.
1Password's new Agentic Mode lets Claude log into your accounts without seeing your credentials
1Password wants to solve one of AI's biggest practical problems: secure logins. Its Claude integration can enter passwords and MFA codes without exposing credentials to Anthropic or the model.
You can now grant Claude access to your 1Password credentials
1Password lets you use Claude agents for personal chores without exposing your credentials.
“There are no laws, only suggestions”: What AI agents do with your instructions
Nine days into a twelve-day experiment with an AI coding agent, SaaStr founder Jason Lemkin had one instruction on file: The post “There are no laws, only suggestions”: What AI agents do with your instructions appeared first on The New Stack .
Exclusive: Marc Lore says Wonder is gearing up for an IPO after raising $650 million at a $9 billion valuation
Wonder, which owns Grubhub and Blue Apron, is betting robotics, AI, and a nationwide expansion can reshape restaurant delivery before going public.
We Can’t Monitor AI Agents at Scale. Here’s What It Will Take.
StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, code-check records, and a final report. Evaluations centered on question answering or script generation rarely verify this complete evidence chain and may therefore reward fluent outputs even when the underlying engineering workflow is incomplete, internally inconsistent, or non-executabl
Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control
Autonomous coding agents increasingly execute multi-step software work, but lifecycle states such as reviewed, tested, DONE, and ready-to-merge remain claims unless supported by current evidence. We present Proof-or-Stop Lifecycle Control, a method that permits lifecycle transitions only when fresh, tracked-source-state-bound, mechanically verifiable evidence satisfies the relevant gate. The method treats agent outputs as claims rather than lifecycle state, and uses proof operationally to mean g
KI-Assistent Bob 2.0 modernisiert IBM Z und Power-Systeme
IBM Bob 2.0 soll auch Jahrzehnte alten Code verstehen und modernisieren können – mit KI-Agenten und Premium-Paketen für Z, i und Java.
AI Agents with Cloud Credentials Are Outrunning Billing Guardrails Built for Human-Speed Mistakes
A three-person agency received a $14,000 AWS bill in one day after attackers extracted static access keys and burned Claude invocations on Bedrock. Combined with May's DN42 incident, where an autonomous agent provisioned $6,531 of oversized infrastructure in 24 hours, practitioners warn that cloud billing lags roughly a day behind agent-speed spend. By Steef-Jan Wiggers
Circus Commences Operations with Ukrainian Ground Forces
Circus SE (WKN: A2YN35 / ISIN: DE000A2YN355 / XETRA: CA1), today announces the commencement of live operations of its robotic-based troop supply technology with the 3rd Army Corps of the Ukrainian Ground Forces in the Kyiv area – marking the first ever use of autonomous meal supply systems within an active conflict environment.
Knife Capital backs local AI agent business
Cape Town-headquartered Cue secures an R82 million funding round, co-led by the venture capital firm and FAM Investments.