Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
From maritime trade to commercial nuclear power, insurance has been the enabler of major economic and technological developments by pricing risk, limiting downside, and spreading best practices. The emerging AI agent economy, projected to handle trillions of dollars in transactions by 2030, looks to be the next such development. Yet insurers' exposure to AI agent risk currently sits largely unpriced across existing insurance lines; between this silent coverage and growing exclusions, coverage is
Turing Award winner Rich Sutton founds Oak Lab to build AI agents that learn on their own
Richard Sutton, 2024 Turing Award winner and co-founder of modern reinforcement learning, has launched a new startup called Oak Lab in Toronto. He calls current deep learning methods "weak and inefficient" and wants to build AI agents that learn continuously from their environment. The article Turing Award winner Rich Sutton founds Oak Lab to build AI agents that learn on their own appeared first on The Decoder .
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state mutations). An automated scenario generation pipeline produces
Narmi releases AI to streamline account opening for communitty banks and credit unions
Narmi, a leading digital banking platform provider for banks and credit unions, today announced the upcoming launch of AI Decision Assist, a new agentic AI capability designed to help financial institutions automate and accelerate account opening reviews while still maintaining control, transparency, and compliance.
Forgetting Our Way to Shared Meaning: Effects of Forgetting on Conceptual Alignment in a Non-Partnership Coordination Game
Shared meaning in language requires people to learn and agree on categories. We ask how characteristics of agents' memories change the emergence and evolution of shared meaning. Without a coordination game, models of conceptual semantics cannot explain how shared meaning emerges and changes in groups of people; however, existing games assume that players share payoffs in a partnership setting. We model conceptual alignment as a non-partnership game and illustrate differences in actual and percei
Paradoxes of Game Theoretic Equilibria and Price of Anarchy
For decades, static solution concepts (Nash, Correlated, and Coarse Correlated Equilibria) and the Price of Anarchy (PoA) have formed the bedrock of algorithmic game theory, with no-regret learning proving fast convergence to such game-theoretic equilibria. We show that reducing multi-agent learning to static equilibrium and black-box regret analysis obscures underlying dynamic disequilibrium and game theoretic bounds. First, interior Nash equilibria lack $C^1$ vector field information, meaning
FIRST Global and Experiential Bring Agentic AI Learning Experience to 190+ Countries, Advancing Robotics Education
FIRST Global Joins the UN ITU AI Skills Coalition; FIRST Global and XRP Kits Offer New Agentic AI module powered by FYI.AI For Student Robotics Teams Worldwide GENEVA, Switzerland — July 8, 2026 — At the AI for Good Global Summit hosted... The post FIRST Global and Experiential Bring Agentic AI Learning Experience to 190+ Countries, Advancing Robotics Education appeared first on AI for Good .
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive information. Existing scam-detection systems mainly focus on isolated messages, which renders them inadequate against this evolving threat. This paper extends single-message phishing detection and presents an explainable agentic system for detecting sophisticated
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive information. Existing scam-detection systems mainly focus on isolated messages, which renders them inadequate against this evolving threat. This paper extends single-message phishing detection and presents an explainable agentic system for detecting sophisticated
Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools. Existing approaches mainly optimize attack success and preserve artifacts such as benchmarks, payloads, or attack programs, which record where attacks succeed but not the enabling conditions behind unsafe agent behavior. We study automated red-teaming for productio
Requirement-Driven Design of Whole-Body Social Tactile Sensing via Virtual Human-Robot Interaction
Tactile sensing for social-physical human-robot interaction (spHRI) is designed in a hardware-driven manner, where predefined sensor configurations constrain coverage, spatial resolution, and the range of recognizable gestures. We propose a requirement-driven framework that derives sensing requirements, specifically spatial resolution and placement, directly from interaction data. Using a VR-based platform with haptic feedback, we collected high-resolution whole-body contact distributions across
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences. However, progress remains fragmented: models use incompatible action spaces and prediction targets, datasets and tasks follow different conventions, and runtime systems e
Now, defenders are embracing the prompt injection, too
"Context bombing" tricks hacking agents into shutting down before they can do harm.
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model
Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling
Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality. Cumulative prospect theory (CPT) has been widely recognized as an effective framework for characterizing such behavioral patterns. However, its large-scale application, particularly in simulation and agent-based modeling, critically depends on specifying individual-level CPT parameters, which remain a major bottleneck. Conventional approaches typically rely
Auditing the Risk Claims of Distributional Reinforcement Learning
Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominan
How DoorDash Built an AI Shopping Assistant That Doesn’t Rely on the LLM Alone
DoorDash details the architecture behind Ask DoorDash, its AI-powered conversational shopping assistant, combining LLMs, specialized AI agents, MCP-based tooling, and an intelligence layer with persistent consumer memory and live backend data. Early results show up to 24% higher checkout conversion, 17% larger baskets, and improved intent accuracy using memory-backed sessions. By Leela Kumili
ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions
As robots become increasingly integrated into human environments, their ability to detect and respond to errors remains critical for maintaining user trust and interaction quality. While recent advances in machine learning have improved error detection capabilities, most approaches are limited to specific contexts, controlled settings, or pre-extracted features, limiting their generalizability and applicability to real-world conditions. To address this challenge, the third edition of the ERR@HRI
Heuristic Learning for Active Flow Control Using Coding Agents
Active flow control involves nonlinear dynamics, partial observations, and computationally expensive simulations, making controller design particularly challenging. Deep reinforcement learning (DRL) has emerged as a powerful framework for such problems, but its success typically relies on large numbers of simulator interactions and produces neural-network policies whose decision process often remains difficult to interpret. In this work, we investigate a different paradigm: instead of optimizing
PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing
Researchers organize the papers they collect into personal folder hierarchies in reference managers, and route each new paper into the folder where it belongs. This task differs from standard hierarchical text classification. A user's folder hierarchy is not a fixed, shared taxonomy but a private and evolving folksonomy whose folder meanings may be topical, shorthand, venue-based, or process-oriented, and are often defined by the papers already stored inside them. We formalize this setting as pe
Technical Report on the CVPR 2026@AdvML Workshop Challenge
Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images an
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstra
Agentic Skill Optimization over Lie Algebroids
Agentic systems increasingly improve themselves by editing skills: prompts, rubrics, plans, tool contracts, examples, validators, and traces. Skill edits are not independent coordinates in a vector space: they are local repairs to structured artifacts whose effects are observed only after rollout, validation, and critique. Distinct edits can have the same immediate visible effect while differing in routing context, template state, guardrail scope, or future composability. The order of edits can
Empowering India’s next generation of innovators with ATL Saathi
Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.
Dignity Without Autonomy
In Prajwala v. Union of India, the Supreme Court held that victims of trafficking for commercial sexual exploitation have a right to rehabilitation under Article 23 read with the right to dignity under Article 21. While the judgment has been celebrated for its three-dimensional dignity framework, it is a missed opportunity to articulate a constitutional basis for protecting the rights of sex workers. The Court's dignity framework – calibrated against objectification in trafficking – is insuffici
Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA
Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web pages, and computation results. Existing agentic multimodal systems often leave evidence in scratchpads, tool trajectories, or free-form histories, making it difficult to track what has been grounded, what remains missing, and when the evidence is sufficient to answer. We propose Omni-Decision, a training-free evidence-state system that turns omni-modal QA i
Senator Warner Makes a First Foray into Agentic AI Regulation
ChinAI #366: Most Companion Robots Die by Day 30
Greetings from a world where…
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model
Health Insurance Marketplaces: CMS Needs Stronger Controls to Prevent Unauthorized Actions by Agents and Brokers
What GAO Found Millions of consumers rely on the assistance of health insurance agents and brokers to purchase health insurance plans through federal and state Marketplaces established by the Patient Protection and Affordable Care Act. The federal Marketplace is maintained by the Centers for Medicare & Medicaid Services (CMS). To assist consumers in the federal Marketplace, agents and brokers must be licensed to sell health plans and be registered with the Marketplace, among other things. CMS co
Health Insurance Marketplaces: CMS Needs Stronger Controls to Prevent Unauthorized Actions by Agents and Brokers
What GAO Found Millions of consumers rely on the assistance of health insurance agents and brokers to purchase health insurance plans through federal and state Marketplaces established by the Patient Protection and Affordable Care Act. The federal Marketplace is maintained by the Centers for Medicare & Medicaid Services (CMS). To assist consumers in the federal Marketplace, agents and brokers must be licensed to sell health plans and be registered with the Marketplace, among other things. CMS co
Exclusive: 34 CEOs on what thrills and terrifies them about agentic AI
When businesses leaders think about AI right now, they’re thinking about how agentic tools will change the very nature of their work. “We’ve crossed a line,” says Varun Krishna, CEO of the fintech giant Rocket Companies . “AI is no longer just creating. It is thinking, deciding and acting. That changes everything, from client interaction to security.” As agentic tools get more sophisticated, they “expose how many organizations a
Building a Foundation Stack for General-Purpose Robots
This article is brought to you by X Square Robot . Large language models gave artificial intelligence a working recipe. Pretrain a large model on broad data, and general capability follows. Robotics has no such recipe. Robotics systems have long been assembled from separate perception, planning, and control parts that rarely add up to intelligence a robot can carry from one task to another, or one machine to another. The central problem in embodied AI is to find the equivalent recipe, and the fi
Chinese internet firms sign AI agent data protection pact
The China Internet Association released a self-regulatory pact on personal information protection for AI agents at a forum in Beijing, with Baidu, Tencent, Alibaba, Volcengine, and 27 other internet companies among the first signatories. The pact is aimed at standardizing how AI agents collect, process, and use personal data as agent-based services spread across internet […]
China sets 2030 target for next-generation internet infrastructure
China’s Ministry of Industry and Information Technology and three other agencies issued guidelines on July 13 to upgrade the country’s internet basic resources, targeting “systematic breakthroughs” by 2030 and a more advanced national internet infrastructure by 2035. The document calls for research into agent-to-agent networks, satellite internet, digital identity infrastructure, IPv6 upgrades, and the integration […]
Ant Group unveils AI safety models for agents and multimodal systems
Ant Group’s AI Safety Lab has open-sourced SingGuard-NSFA, a safety guardrail model for autonomous agents, and disclosed details of SingGuard, a multimodal safety model. SingGuard-NSFA is designed to detect risks such as prompt injection, sensitive data theft, malicious code execution, resource abuse, and permission misuse before agents take action. The model covers seven major risk […]
Towards Predictive, Aligned, and Scalable Robot Learning
Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding future possibilities and providing a unified substrate for cross-modal alignment. This formulation enables predictive reasoning akin to
Vietnam won two second prizes at the global finals of the Robotics for Good 2026 competition
According to information from the STEM Education Promotion Alliance (SEPA) on July 11th, the two Vietnamese teams excellently won second place in the Junior and Senior categories of the Robotics for Good Youth Challenge 2026 Global Finals. The post Vietnam won two second prizes at the global finals of the Robotics for Good 2026 competition appeared first on AI for Good .
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments. This work departs from the capability first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination resistant benchmarks, and interactive training e
Fast-tracking AI: Linking data, governance, easy access
Artificial intelligence agents can both read data and write actions, meaning security and governance are critically important.