Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
Google's Genkit Ships Agents API with Detached Turns and Human-in-the-Loop for TypeScript and Go
Google released the Genkit Agents API in preview for TypeScript and Go. The open-source framework packages message history, tool loops, streaming, and state persistence behind a single chat() interface. Detached turns let agents work after clients disconnect. Interruptible tools provide human-in-the-loop control with anti-forgery validation on resume. By Steef-Jan Wiggers
Community leaders, organizers rally to demand accountability after ICE shooting in Maine
Community leaders and organizers rallied to demand accountability after an ICE agent fatally shot a motorist in Biddeford, Maine on Monday.
What makes CIOs trust an AI agent? Thira bets it’s not the model.
Sunny Gupta spent a decade and a half building Apptio into the system of record for enterprise technology spend. His The post What makes CIOs trust an AI agent? Thira bets it’s not the model. appeared first on The New Stack .
How to manage AI investments in the agentic era
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.
Waymo’s July 4 chaos in San Francisco raises new questions about how robotaxis can work at scale
Call it AV déjà vu: Once again, Waymo’s performance during a major disruption is raising questions about whether it is ready to operate at scale. Over the July 4 weekend, Waymo vehicles clogged San Francisco streets near the city’s fireworks celebration. A string of robotaxis, apparently unaware of event-related road closures, worsened the already severe gridlock. The batteries in some of the autonomous vehicles died, requiring them to be towed. At least one Waymo drove straight through a firewo
Robotik: MIT-Roboter fliegt und schwimmt wie ein Papageitaucher
Forscher des MIT entwickeln einen fliegenden und schwimmenden Roboter nach dem Vorbild von tauchenden Seev�geln. ( Roboter , Wissenschaft )
Chinese AI giants, start-ups ready for the agentic web era
The agentic web era has arrived, and Chinese companies are going all in. Humans were overtaken last month by artificial intelligence (AI) agents and AI bots as the majority users of the internet, data from global network company Cloudflare shows. By early July, agents and bots accounted for more than 60 per cent of web traffic. Even Cloudflare CEO Matthew Prince expressed surprise in a post on X, having initially forecast AI agent-driven web use would pass the 50 per cent mark only in late 2027.
Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequently, safety is no longer only about input--output content alignment. It also concerns system behavior and real-world execution outcomes. However, the current literature is fragmented across attack types, applications, and benchmarks. This makes it hard to explain why failures such as prompt injection, tool misuse, and memory poisoning often share t
China’s electric vehicle exports rise 68.7% in H1 2026
China’s electric vehicle exports rose 68.7% year-on-year in the first half of 2026, according to customs data released on July 14. Exports of electric motorcycles and bicycles increased 31.5%, electric locomotives 45.1%, lithium batteries 37.6% and wind turbines 35.6%. The same data showed that exports of AI-integrated intelligent bionic robots exceeded 10,000 units and reached […]
AI lawsuits expose gaps in conventional insurance, says report
Shift from chatbot errors to autonomous ‘agents’ leaves businesses exposed to widening range of legal claims
From Tools to Teacher-Built Teammates: No-Code Pedagogical Plugin Authoring with LearnAdapt Agentic Studio and PedOS 1.1 Lumina
arXiv:2607.09674v1 Announce Type: new Abstract: Teachers and researchers need to adapt educational AI to local goals, but most systems remain difficult to customize or study without coding expertise. We present LearnAdapt Agentic Studio on PedOS 1.1 Lumina, a no-code authoring and governed runtime environment for educational AI plugins. A non-coder describes a desired learning interaction in plain English; the system prepares a previewable plugin artifact, runs safety checks, and supports submis
Robo-Reporters: Evaluating Autonomous AI Agents as Algorithmic Gatekeepers in Computational Journalism
arXiv:2607.10736v1 Announce Type: new Abstract: Artificial intelligence agents increasingly perform journalism tasks autonomously, searching for sources, evaluating credibility, and producing news content with minimal human oversight. Yet research has largely treated AI as a monolithic category, leaving the effects of architectural design unexamined. Drawing on gatekeeping theory, this study presents the first systematic comparison of four agent architectures, monolithic (Claude), chain-based (L
Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems
arXiv:2607.09685v1 Announce Type: cross Abstract: Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions. In adaptive socio-technical systems this assumption may fail: regulatory change can alter incentives, agents can respond strategically, and the mapping from policy variables to aggregate outcomes can change. This paper studies such regime change as a transfer-learning problem in adaptive multi-agent systems.
Intervenability as a Design Requirement for Autonomy and Oversight within Human-Centered AI
arXiv:2607.10322v1 Announce Type: cross Abstract: Based on the literature and several practical examples of possible AI applica-tions, we outline the concept of intervenability. This new phenomenon is not covered by emergency shutdowns, workarounds, or the reconfiguration of automated systems. Intervenability instantiates the principles of control-lability, autonomy, oversight, and keeping humans in the loop in the context of AI. We provide a taxonomy that encompasses a range of possibilities fo
Automated Textbook Auditing with Multi-Agent LLM Systems
arXiv:2607.11276v1 Announce Type: cross Abstract: Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, domain-specific technical correctness, and linguistic quality simultaneously -- a task that general-purpose grammar checkers cannot address. We present \textbf{AI Textbook Auditor}, a modular multi-agent pipeline for automated quality assurance of educational materials across subject domains. The system accepts a
Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning
arXiv:2604.22770v2 Announce Type: replace Abstract: Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. When progression is driven by quiz performance, learners can advance despite persistent gaps in using grammar and vocabulary during interaction. Recent work on LLM-based judging suggests a path toward scoring open-ended conversations, but using interaction evidence to drive progression and review requires scori
Uber’s product chief on hotels, robotaxis, and why the company doesn’t want to be ‘everything for everyone’
Uber Chief Product Officer Sachin Kansal walks TechCrunch through the company's financial services ambitions, its increasingly complicated relationship with Waymo, its new AV Labs data operation, and how AI is starting to show up in ways riders and drivers will actually notice.
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
Proactive agents that anticipate user needs and autonomously execute tasks hold great promise as digital assistants, yet the lack of realistic user simulation frameworks hinders their development. Existing approaches model apps as flat tool-calling APIs, failing to capture the stateful and sequential nature of user interaction in digital environments and making realistic user simulation infeasible. We introduce Proactive Agent Research Environment (Pare), a framework for building and evaluating
AI agents are becoming the enterprise's most privileged users. And most organisations don't know it yet
AI agents are outpacing identity governance.
Hermes agent maker Nous Research in talks for new funding at $1.5B valuation
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
Analysis of Mutual and Referential Human and Robot Gazes in a Collaborative Word Association Game
Robot gaze is a major component of human-robot dialogue coordination. Most studies of gaze in human-robot dialogue focus on face-to-face social conversations, but little is known about gaze in demanding task-focused interactions. In this paper, we investigate how the gaze of a robot game partner affects human visual attention and if humans tend to direct confirmation-seeking gazes towards the robot. In our study, we let participants play a collaborative word association game with a NAO robot act
datasette code-frequency chart on GitHub
datasette code-frequency chart on GitHub Out of curiosity I decided to see if I could find a useful illustration of the impact of coding agents and Opus 4.5 class models on my own output. The best I've found so far is this GitHub chart of frequency of code changes to my Datasette open source project: The big spike in activity at the end aligns with Opus 4.8, GPT-5.5, Fable 5 and GPT-5.6 Sol. Tags: github , ai , datasette , generative-ai , llms , ai-assisted-programming , coding-agents
Tracing Agentic Failure from the Flow of Success
Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based pipelines, which are computationally expensive, or require post-training on failure trajectories with step-level error annotations, which are costly to collect and difficult to scale. We argue that a practical failure attribution model should be lightweight and tr
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a change can be made, a developer or coding agent must identify all code locations that implement the target behavior. This is difficult because production harnesses are large, tightly coupled, and behavio
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception via a "skeptical" reasoning paradigm. Unlike holistic models, ChartCynics decouples perception from verification: a Diagnostic Vision Path captures structural anomalies (e.g., inverted axes) through strategic ROI cropping, while
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We explo
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we int
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluation protocols assess and optimize for predefined goals such as capture-the-flag, remote code execution, exploit reproduction, or trajectory similarity, in simplified or narrow settings. These tools are valuable for measuring bounded capabilities, yet they do not adequately capture the complexity, open
Self-Improvements in Modern Agentic Systems: A Survey
Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains. We offer a system-level framework that represents a modern agent as a configuration coupling a foundation model with an operational scaffold of prompts, memory, tools, and
PalmClaw: A Native On-Device Agent Framework for Mobile Phones
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user
Rethinking the Evaluation of Harness Evolution for Agents
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple tas
From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality
Code review helps maintain software quality before code integration, but it also imposes a substantial workload on human reviewers. As generative artificial intelligence becomes part of software development, code review is shifting from a primarily human review process toward AI-supported review processes in which large language model (LLM) reviewers and AI agent reviewers participate alongside human reviewers. However, we still lack empirical evidence on how this transition affects review effic
ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest. Unfortunately, existing systems either keep their memory inside the model context or in a flat feature store, and organize it around frames rather than around the persistent entities a stream is really about, which confines them to
Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap
One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. The field has since moved faster than anticipated. Multi-agent systems have produced experimentally validated hypotheses, self-driving laboratories have grown more interoperable and orchestrated, reasoning-trained and domain foundation models have raised the capability ceiling, and the Genesis Mission has placed autonomous e
U.S. military uses Corsair maritime drones to attack Iran
The robotic platforms, which are built by Saronic, were recently employed in an attack against Iran for the first time. The post U.S. military uses Corsair maritime drones to attack Iran appeared first on DefenseScoop .
CityBehavEx: A Scalable and Empirically Validated LLM-Assisted Urban Simulation Platform
Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validated against empirical mobility patterns. We present CityBehavEx, an interactive LLM-assisted urban simulation platform that scales to city-size populations, exposes agent behavior for inspection, supports empirical validation, and generates mobility patterns that better match real-world spatial, temporal, and semantic distributions. Instead of inv
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality. Although LLM-as-a-judge methods provide scalable alternatives to human evaluation, production deployment introduces challenges in governance, reproducibility, cost, schema consistency, traceability, and reliability. We present GenAI Evaluation, a governed, configuration-driven pipeline for large-scale evaluation
AI agents create virtual playgrounds to help robots get crucial training data
“SceneSmith” system uses collaborative AI agents to create realistic 3D environments of places like kitchens, hotels, and living rooms, where robots can simulate everyday chores.
Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now support both human and agent-mediated interaction. This paper introduces the agent-ready website, a design framework for enhancing the readability, interpretability, verifiability, and actionability of e-commerce platforms for AI agents. Existing web design, SEO, and ge
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manip