05:42 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

Now, defenders are embracing the prompt injection, too

"Context bombing" tricks hacking agents into shutting down before they can do harm.
Ars Technica 18d ago News Agents & autonomy

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model
arXiv cs.AI 18d ago Research Agents & autonomy

Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling

Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality. Cumulative prospect theory (CPT) has been widely recognized as an effective framework for characterizing such behavioral patterns. However, its large-scale application, particularly in simulation and agent-based modeling, critically depends on specifying individual-level CPT parameters, which remain a major bottleneck. Conventional approaches typically rely
arXiv cs.AI 18d ago Research Agents & autonomy

Auditing the Risk Claims of Distributional Reinforcement Learning

Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominan
arXiv cs.AI 18d ago Research Safety & alignmentAgents & autonomy

How DoorDash Built an AI Shopping Assistant That Doesn’t Rely on the LLM Alone

DoorDash details the architecture behind Ask DoorDash, its AI-powered conversational shopping assistant, combining LLMs, specialized AI agents, MCP-based tooling, and an intelligence layer with persistent consumer memory and live backend data. Early results show up to 24% higher checkout conversion, 17% larger baskets, and improved intent accuracy using memory-backed sessions. By Leela Kumili
InfoQ AI/ML 18d ago News Agents & autonomy

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions

As robots become increasingly integrated into human environments, their ability to detect and respond to errors remains critical for maintaining user trust and interaction quality. While recent advances in machine learning have improved error detection capabilities, most approaches are limited to specific contexts, controlled settings, or pre-extracted features, limiting their generalizability and applicability to real-world conditions. To address this challenge, the third edition of the ERR@HRI
arXiv cs.HC 18d ago Research Agents & autonomyEnvironment

Heuristic Learning for Active Flow Control Using Coding Agents

Active flow control involves nonlinear dynamics, partial observations, and computationally expensive simulations, making controller design particularly challenging. Deep reinforcement learning (DRL) has emerged as a powerful framework for such problems, but its success typically relies on large numbers of simulator interactions and produces neural-network policies whose decision process often remains difficult to interpret. In this work, we investigate a different paradigm: instead of optimizing
arXiv cs.AI 18d ago Research Agents & autonomyFinance, VC & PE

PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing

Researchers organize the papers they collect into personal folder hierarchies in reference managers, and route each new paper into the folder where it belongs. This task differs from standard hierarchical text classification. A user's folder hierarchy is not a fixed, shared taxonomy but a private and evolving folksonomy whose folder meanings may be topical, shorthand, venue-based, or process-oriented, and are often defined by the papers already stored inside them. We formalize this setting as pe
arXiv cs.HC 18d ago Research Agents & autonomy

Technical Report on the CVPR 2026@AdvML Workshop Challenge

Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images an
arXiv cs.AI 18d ago Research Agents & autonomy

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstra
arXiv cs.AI 18d ago Research RegulationAgents & autonomy

Agentic Skill Optimization over Lie Algebroids

Agentic systems increasingly improve themselves by editing skills: prompts, rubrics, plans, tool contracts, examples, validators, and traces. Skill edits are not independent coordinates in a vector space: they are local repairs to structured artifacts whose effects are observed only after rollout, validation, and critique. Distinct edits can have the same immediate visible effect while differing in routing context, template state, guardrail scope, or future composability. The order of edits can
arXiv cs.AI 18d ago Research Safety & alignmentAgents & autonomy

Empowering India’s next generation of innovators with ATL Saathi

Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.
Google DeepMind 18d ago Field notes Agents & autonomy

Dignity Without Autonomy

In Prajwala v. Union of India, the Supreme Court held that victims of trafficking for commercial sexual exploitation have a right to rehabilitation under Article 23 read with the right to dignity under Article 21. While the judgment has been celebrated for its three-dimensional dignity framework, it is a missed opportunity to articulate a constitutional basis for protecting the rights of sex workers. The Court's dignity framework – calibrated against objectification in trafficking – is insuffici
Verfassungsblog (EU law incl AI) 18d ago Field notes Agents & autonomy

Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA

Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web pages, and computation results. Existing agentic multimodal systems often leave evidence in scratchpads, tool trajectories, or free-form histories, making it difficult to track what has been grounded, what remains missing, and when the evidence is sufficient to answer. We propose Omni-Decision, a training-free evidence-state system that turns omni-modal QA i
arXiv cs.AI 18d ago Research Agents & autonomy

Senator Warner Makes a First Foray into Agentic AI Regulation

Tech Policy Press 18d ago News RegulationAgents & autonomy

ChinAI #366: Most Companion Robots Die by Day 30

Greetings from a world where…
ChinAI 18d ago Field notes Agents & autonomy

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model
HuggingFace Daily Papers 18d ago Research Agents & autonomy

Health Insurance Marketplaces: CMS Needs Stronger Controls to Prevent Unauthorized Actions by Agents and Brokers

What GAO Found Millions of consumers rely on the assistance of health insurance agents and brokers to purchase health insurance plans through federal and state Marketplaces established by the Patient Protection and Affordable Care Act. The federal Marketplace is maintained by the Centers for Medicare & Medicaid Services (CMS). To assist consumers in the federal Marketplace, agents and brokers must be licensed to sell health plans and be registered with the Marketplace, among other things. CMS co
US GAO Reports 18d ago Policy HealthcareAgents & autonomy

Health Insurance Marketplaces: CMS Needs Stronger Controls to Prevent Unauthorized Actions by Agents and Brokers

What GAO Found Millions of consumers rely on the assistance of health insurance agents and brokers to purchase health insurance plans through federal and state Marketplaces established by the Patient Protection and Affordable Care Act. The federal Marketplace is maintained by the Centers for Medicare & Medicaid Services (CMS). To assist consumers in the federal Marketplace, agents and brokers must be licensed to sell health plans and be registered with the Marketplace, among other things. CMS co
US GAO Reports 18d ago Policy HealthcareAgents & autonomy

Exclusive: 34 CEOs on what thrills and terrifies them about agentic AI

When businesses leaders think about AI right now, they’re thinking about how agentic tools will change the very nature of their work. “We’ve crossed a line,” says  Varun Krishna, CEO of the fintech giant  Rocket Companies . “AI is no longer just creating. It is thinking, deciding and acting. That changes everything, from client interaction to security.”  As agentic tools get more sophisticated, they “expose how many organizations a
Fast Company Tech 18d ago News Agents & autonomyFinance, VC & PE

Building a Foundation Stack for General-Purpose Robots

This article is brought to you by X Square Robot . Large language models gave artificial intelligence a working recipe. Pretrain a large model on broad data, and general capability follows. Robotics has no such recipe. Robotics systems have long been assembled from separate perception, planning, and control parts that rarely add up to intelligence a robot can carry from one task to another, or one machine to another. The central problem in embodied AI is to find the equivalent recipe, and the fi
IEEE Spectrum 18d ago News Agents & autonomy

Chinese internet firms sign AI agent data protection pact

The China Internet Association released a self-regulatory pact on personal information protection for AI agents at a forum in Beijing, with Baidu, Tencent, Alibaba, Volcengine, and 27 other internet companies among the first signatories. The pact is aimed at standardizing how AI agents collect, process, and use personal data as agent-based services spread across internet […]
TechNode (CN) 18d ago News RegulationPrivacy

China sets 2030 target for next-generation internet infrastructure

China’s Ministry of Industry and Information Technology and three other agencies issued guidelines on July 13 to upgrade the country’s internet basic resources, targeting “systematic breakthroughs” by 2030 and a more advanced national internet infrastructure by 2035. The document calls for research into agent-to-agent networks, satellite internet, digital identity infrastructure, IPv6 upgrades, and the integration […]
TechNode (CN) 18d ago News Agents & autonomy

Ant Group unveils AI safety models for agents and multimodal systems

Ant Group’s AI Safety Lab has open-sourced SingGuard-NSFA, a safety guardrail model for autonomous agents, and disclosed details of SingGuard, a multimodal safety model. SingGuard-NSFA is designed to detect risks such as prompt injection, sensitive data theft, malicious code execution, resource abuse, and permission misuse before agents take action. The model covers seven major risk […]
TechNode (CN) 18d ago News Safety & alignmentAgents & autonomy

Towards Predictive, Aligned, and Scalable Robot Learning

Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding future possibilities and providing a unified substrate for cross-modal alignment. This formulation enables predictive reasoning akin to
arXiv 18d ago Research Safety & alignmentAgents & autonomy

Vietnam won two second prizes at the global finals of the Robotics for Good 2026 competition

According to information from the STEM Education Promotion Alliance (SEPA) on July 11th, the two Vietnamese teams excellently won second place in the Junior and Senior categories of the Robotics for Good Youth Challenge 2026 Global Finals. The post Vietnam won two second prizes at the global finals of the Robotics for Good 2026 competition appeared first on AI for Good .
AI for Good (ITU) 18d ago Field notes Children & educationAgents & autonomy

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments. This work departs from the capability first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination resistant benchmarks, and interactive training e
arXiv 18d ago Research HealthcareAgents & autonomy

Fast-tracking AI: Linking data, governance, easy access

Artificial intelligence agents can both read data and write actions, meaning security and governance are critically important.
ITWeb (ZA) 18d ago News RegulationPrivacy

NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management

Large language models (LLMs) based agents are beginning to participate in portfolio construction and market analysis, where decisions must be justified under evolving information and risk constraints. Current assessment practice, however, remains poorly aligned with this setting: many studies rely on static examinations or report only terminal portfolio returns, while the intermediate evidence, analyst judgments, and execution steps that produced those returns stay largely invisible. We introduc
arXiv 18d ago Research PrivacyAgents & autonomy

A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery

The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands of options simultaneously. This leads to decision-space explosion, context window saturation, and degraded routing accuracy. To address these limitations, this paper presents a hierarchical, skill-based architecture for agentic orchestration. Capabilities are or
arXiv 18d ago Research Agents & autonomy

NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study

Agentic research systems are emerging as a new paradigm for coordinating scientific workflows beyond isolated model inference, code generation, or statistical analysis. However, deployment in institutional biomedical environments requires governed mechanisms for research planning, data access, workflow orchestration, evidence tracking, reproducibility, and human oversight. We present NVAITC AI Scientist (NAIS), a governed end-to-end agentic research system designed to support domain-general scie
arXiv 19d ago Research PrivacyAgents & autonomy

Multi-Agent LLMs Fail to Explore Each Other

Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one another. We show that modern LLM agents fail to do so, often exhibiting myopic and polarized interaction patterns that lead to suboptimal coordination and increased regret. We formalize this challenge as the Multi-Agent Exploration problem, modeling it as a partially observable stochastic game (POSG) problem in w
HuggingFace Daily Papers 19d ago Research Agents & autonomy

The Agentic Age Needs A Cognitive Operating Model

Last October, I published a blog proposing a different mental model for AI agents: Treat them as cognitive skills and products, not as digital employees. That framing has since resonated strongly with Forrester clients, particularly technology leaders building agentic capabilities inside the enterprise. But the concept of a cognitive skill in that blog was deliberately loose. […]
Forrester AI blog 19d ago Field notes Agents & autonomy

QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics

Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call for a data-agent architecture that treats semantics, methodology, execution, and evolution as first-class system concerns. To this end, we introduce QwenPaw-Data, an agentic data system designed for enterprise intelligent data analysis. QwenPaw-Dat
arXiv 19d ago Research Agents & autonomyEnvironment

China’s massive AI rollout - podcast

Senior China correspondent Amy Hawkins on China’s embrace of AI, from medical avatars to food delivery drones and state surveillance While the spread of AI has been met perhaps with a lot of scepticism in the west, China has fully embraced the technology, explains Amy Hawkins , from millions of users talking to AI doctors, to the use of intelligent robots in factories, and drones delivering food on the Great Wall of China. AI has also been eagerly taken up by the state, not least in the opportun
The Guardian 19d ago News PrivacyHealthcare

Directly Responsible Individuals (DRI)

Directly Responsible Individuals (DRI) I went looking for a definition of "Directly Responsible Individuals" and the best I found was in the GitLab handbook. Apparently the term originated at Apple, where it's used to describe the person who is "ultimately accountable for the success or failure of a specific project, initiative, or activity". I've been thinking about this term recently in the context of LLM-powered agents and how they fit into human organizations. I don't think an agent should e
Simon Willisons Weblog 19d ago Field notes Agents & autonomyTransparency

Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution

LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods attempt to mitigate this limitation through pre-repair repository exploration; however, their fix-driven strategies explore repositories without identifying the agent's knowledge gaps, often yielding imprecise context that fails to bridge the underlying understanding deficit. In this paper, we propo
HuggingFace Daily Papers 19d ago Research Agents & autonomy

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstra
HuggingFace Daily Papers 19d ago Research RegulationAgents & autonomy

A Vocabulary for Multi-Agent Automated Research Systems

We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies 1) who the agents are, 2) what operations are available in the system, 3) who may invoke them, 4) how agents communicate, 5) what information is visible within and across runs, 6) how the next action is chosen, 7) how a run begins, and 8) how outputs are evaluated. A trajectory records one run from the input task to the retur
HuggingFace Daily Papers 19d ago Research Agents & autonomy

LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans

AI agents are evolving from answer engines into persistent teams that use tools, delegate work, learn from experience, and modify the artifacts that shape their future behavior. The defining question for deployment is no longer merely what agents can do, but who controls what they are allowed to become. We introduce logos, a pluggable layer for self-evolution and governance that strengthens existing multiagent frameworks rather than replacing them. logos compiles heterogeneous multimodal inputs,
arXiv 19d ago Research RegulationAgents & autonomy
← Newer Older →