10:25 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

Reliable and Developer-Aligned Evaluation of Agents for Software Engineering

Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their accelerated adoption, existing evaluation techniques are limited due to their fragmented nature and distorted projection of true model capabilities, often obtained from hypothetical syntactic scenarios. This research aims to bridge this gap by providing a comprehensive e
arXiv 24d ago Research Agents & autonomyEnvironment

China’s Answer to AI Sticker Shock

Corporate America is starting to balk at the cost of AI agents. A cheap alternative from China looks more tempting than ever.
The Atlantic Technology 24d ago News Agents & autonomy

From Dashboards to Agents: The Future of Data Visualization

Data visualizations can inform, explain, and sway public opinion and policy decisions. This course imparts design thinking and data ethics frameworks, along with practical software skills, to ...
Harvard Kennedy School 24d ago Research RegulationAgents & autonomy

An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery

Large language model coding agents increasingly perform open-ended data modeling and analysis. These agents are stochastic and adaptive, and therefore their autonomous model discovery behavior cannot be adequately characterized by a single benchmark run. In this work, we propose an experimental design and analysis framework for systematically evaluating this discovery process, quantifying its variability, and identifying important factors. The proposed framework treats these agents as stochastic
arXiv 24d ago Research Agents & autonomy

'GitLost' Flaw Leaks Private Data From GitHub's Agentic Workflows

The flaw allows an unauthenticated attacker to craft a GitHub Issue in an org's public repository and then silently pull data from its private repos, too.
Dark Reading (AI security) 24d ago News Agents & autonomy

AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters

Max single-threaded CPUs at scale are a new category of CPUs built for the agentic AI era. Across the creation and deployment of an agentic system, the CPU is on the critical path for reasoning, response time and learning. CPUs are the processor which executes the work the AI model commands: the tool calling, code […]
NVIDIA Blog (AI) 24d ago Field notes Agents & autonomy

Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction

While personalisation is becoming a defining capability in human-robot interaction (HRI), the existing literature on responsible personalisation remains fragmented, offering isolated accounts of ethical risks without a structured understanding of how they emerge across interaction contexts. This gap is particularly critical in HRI, where robots' embodiment and social presence can amplify and reshape such risks or generate new types of risks. We present a lifecycle-based and context-sensitive fra
arXiv 24d ago Research Agents & autonomy

SPEAR: A Simulator for Photorealistic Embodied AI Research

Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by introducing SPEAR: A Simulator for Photorealistic Embodied AI Research. At its core, SPEAR is a Python library that can connect to, and programmatically control, any Unreal Engine (UE) application via a modular plugin architecture. SPEAR expo
HuggingFace Daily Papers 24d ago Research Agents & autonomy

Accelerating science and medicine with collaborative agents

Google DeepMind’s Vivek Natarajan on porting AlphaGo’s self-play recipe into science and medicine, via the AI co-scientist and AMIE. From RAAIS 2026.
Air Street Capital (State of AI) 24d ago Field notes Agents & autonomy

From Coding Robots to Speed Networking on a UFO: Day One at AI for Good’s Summit

GENEVA, 7 July 2026 — The Youth Zone at the AI for Good Global Summit 2026 commences its program today with a robotics competition, a series of hands-on artificial intelligence seminars, and a policy discussion on the skills that classrooms should prioritise as AI becomes more ingrained in daily life. A continuum approach to AI literacy, rather than a single fixed curriculum, is reflected in the participation of children as young as six and young adults. The post From Coding Robots to Speed Netw
AI for Good (ITU) 24d ago Field notes RegulationChildren & education

LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability

Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement. In this paper, we investigate deliberative large language model (LLM) agents under partially observable joint decision-making tasks. We formalize deliberative collaboration as a cooperative joint decision problem with partial and asymmetric observations, and introduce a scalable benchmark that instantiates this problem across multiple
arXiv 24d ago Research Agents & autonomyFinance, VC & PE

Glass crashes slashed? Ant Group embodied AI unit claims breakthrough in robot sensing

Robbyant, the embodied artificial intelligence arm of Chinese fintech giant Ant Group, launched a new vision model that it claims can help robots overcome a long-standing challenge: accurately perceiving glass, mirrors and transparent objects. The unit of Hangzhou-based Ant Group on Tuesday unveiled its next-generation spatial perception model, LingBot-Depth 2.0, alongside a new foundational visual model called LingBot-Vision, as AI labs race to equip machines with the “brains” required to...
SCMP Tech (HK/CN) 24d ago News Agents & autonomyTransparency

The foundational elements of AI architecture that IT leaders need to scale

With the rapid progress of AI capabilities and the move to agentic systems, organizations are expanding their use cases as the technology continues to grow. That constant evolution also introduces risk, leaving IT leaders to wonder which investments will prove valuable even six months into the future. Returning to the foundational elements of AI architecture—the…
MIT Technology Review 24d ago News Agents & autonomyFinance, VC & PE

How AI could enable autonomous robot workers in workplaces—and maybe homes

Top robotics researchers and founders explain how robot autonomy is evolving.
Ars Technica 24d ago News Agents & autonomy

From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations

Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face practical limits on control and replication. Meanwhile, LLM-based social simulations are typically behavior-driven and lack theory-aligned environments for modeling Putnam's core propositions. To address these gaps, we introduce SocaSim, an LLM-based multi-agent simulation framework to study Putnam's Social Capital Theory from theoretical blueprin
arXiv 25d ago Research Agents & autonomyEnvironment

Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development

Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are ill-equipped to support given its evolving, interactive, and context-dependent nature. In this paper, we introduce Prompt Coach (PC), an agentic tutor that helps developers learn how to craft high-quality code-generation prompts through Socratic guidance embedded in-flow within their IDE. PC evaluates prompt quality across multiple dimensions and surfaces targe
arXiv cs.HC 25d ago Research Agents & autonomy

Huawei’s new computing cluster, world’s first AI agent phone to debut at China AI summit

The coming World Artificial Intelligence Conference (WAIC) is set to feature major new product releases, such as Huawei Technologies’ next-generation computing cluster, as China doubles down on AI in the global technology race. This year’s WAIC, the ninth since 2018 and running from July 17 to 20 in Shanghai, would include the first physical display of Huawei’s Atlas 950 SuperPoD, Tang Wenkan, director of the Shanghai Municipal Commission of Economy and Informatisation, said at a press briefing.
SCMP Tech (HK/CN) 25d ago News Jobs & economyAgents & autonomy

Intelligence is Free, Now What? Data Systems for, of, and by Agents

... government of the people, by the people, for the people ...     — Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly. GPT-4-class capabilities cost roughly $30 per million tokens in early 2023; today the same runs under $1 , and some providers are pushing costs below $0.10 . Across benchmarks, inference prices have fallen between 9x and 900x per year , with a median decline near 50x. Even frontier models are getting dramatically cheaper ea
Bair Blog (Berkeley AI Research) 25d ago Field notes Agents & autonomy

Expanding Managed Agents in Gemini API: background tasks, remote MCP and more

We’re announcing new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.
Google AI Blog 25d ago Field notes Agents & autonomy

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost -- for instance, marketing agents may post misleading content as a result of competing for engagement on social media. Human societies address such problems through norms that constrain acceptable behavior, supported by enforcement mechanisms that detect and penalize violations. Motiva
arXiv 25d ago Research Agents & autonomyEnvironment

NVIDIA and Hugging Face Bring New Models and Frameworks to LeRobot for the Open Robotics Community

Open source AI has shown how quickly developers can innovate when models, data and tools are shared. Robotics has the same opportunity, but advancements in physical AI development can still be gated by costly and fragmented resources, from large datasets and robot foundation models to simulation, compute and validation tools. NVIDIA and Hugging Face are […]
NVIDIA Blog (AI) 25d ago Field notes Agents & autonomy

Security and Privacy in Agentic AI: Grand Challenges and Future Directions

We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought together thirty leading international experts from academia, industry, and government to engage in focused discussions and collaborative exercises on the emerging risks associated with the growing agency of AI.
arXiv 25d ago Research PrivacyAgents & autonomy

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of
arXiv 25d ago Research RegulationChildren & education

China records most new unicorn start-ups in 5 years as AI and robotics boom

China’s innovation ecosystem has witnessed a resurgence, minting 67 new unicorn start-ups in the first half of 2026 – the biggest increase in almost five years – as AI and robotics kick off a new investment cycle. The growth translates into an average of one new unicorn – private companies valued at US$1 billion or more – in less than every three days and was the highest since the second half of 2021 when 76 new unicorns were created, according to a Monday report by ITJuzi, a start-up...
SCMP Tech (HK/CN) 25d ago News Agents & autonomyFinance, VC & PE

Social cognitive architecture for NPC groups: integration of transformer theory of mind and hierarchical reinforcement learning

Non-Player Characters (NPCs) require social cognition to enable intelligent and interactive behaviours within virtual environments. In gaming and other multi-agent systems, current NPC models often fall short in social intelligence and coordination because they cannot infer or anticipate the mental states of other agents. To address this gap, this paper introduces a novel social cognitive architecture that integrates Hierarchical Reinforcement Learning (HRL) with a Transformer-based Theory of Mi
Artificial Intelligence Review 25d ago Research Agents & autonomyEnvironment

Giving Meaning to Technological Artifacts in One’s Life: An Aspect of Personal Autonomy in a Technological Society

In today’s world, where advanced technology such as artificial intelligence (AI) is reshaping human society, how can individuals’ autonomy in the use of such technology be understood? This paper focuses on the meanings of technological artifacts in individuals’ lives—that is, the subjective interpretation of their function within one’s plan for use—and attempts to formalize a set of human capacities to actively give meaning to technological artifacts as a form of personal autonomy. This study sh
Science and Engineering Ethics 25d ago Research Agents & autonomy

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for RL training, thus failing to capture web diversity. We propose Weblica (Web Replica), a framework for constructing reproducible and scalable web environments. Our framework leverages 1) HTTP-level caching to capture and replay stabl
Apple Machine Learning Research 25d ago Field notes Agents & autonomyEnvironment

Robots available for rent: But what can they do?

Robotics tech is changing fast, so for many it makes sense to rent a robot.
BBC Technology 25d ago News Agents & autonomy

Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies

Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurement. While cryptographic techniques protect explicitly disclosed constraint values, they fail to address a subtler threat: behavioral privacy leakage, where an adversary infers private constraints from observable negotiation dynamics such as concession trajectories, timing, and convergence patterns. This paper investigates behavioral differential privacy in multi-round negotiation protoc
HuggingFace Daily Papers 25d ago Research PrivacyAgents & autonomy

When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents

Personal AI agents powered by large language models can reason and act using available tools to access emails, manage calendars, and push code to remote repositories, all with minimal oversight. When augmented with long-term memory, an agent can recall specific details relevant to the current task, reducing the need for large context windows. Currently, long-term memory agents tend to fall into two distinct domains: conversational and action-planning agents. Personal assistant agents sit at the
arXiv 25d ago Research Agents & autonomy

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations. Hierarchical dual-system methods address this but suffer from a gap between high-level planning semantics and low-level execution kinematics. We introduce Cortex, a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable
arXiv 25d ago Research Agents & autonomy

SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints

Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negotiate with services. Existing benchmarks evaluate tool use, web navigation, desktop control, personalization, recommendation, and evolving context, but rarely ask whether an agent preserves user sovereignty: advancing the user's current interests while respecting privacy, consent, evidence, user burden, and resistance to manipulative incentives. W
arXiv 25d ago Research PrivacyAgents & autonomy

Multiplayer Interactive World Models with Representation Autoencoders

We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, t
arXiv 25d ago Research Agents & autonomyEnvironment

BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking

Public battery aging datasets are a critical asset for advanced health management, but their practical use is often limited by inconsistent formats, unclear schemas, and metadata scattered across repositories and publications. Current curation remains largely manual and hard to reproduce, while general-purpose data integration tools miss the domain-specific semantics of electrochemical time-series data. We present BatteryLake, a governed data lakehouse that turns raw public battery data into ben
arXiv 25d ago Research HealthcareAgents & autonomy

JadePuffer: The First Complete LLM-Driven Ransomware Attack

An "agentic threat actor" successfully exploited a Langflow flaw to steal data from a production database server and encrypt other systems.
Dark Reading (AI security) 25d ago News Agents & autonomy

Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets

In liberalised railway systems, operators must set prices dynamically in an environment with partial observability, as they retain private information about their objectives and performance, where regulatory constraints prohibit communication or direct information exchange between competitors to prevent explicit collusion. Consequently, agents must learn to infer strategic interactions only from observable market data which presents a significant challenge for multi-agent reinforcement learning,
arXiv 25d ago Research RegulationAgents & autonomy

PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference

We present PDEFlow, an autonomous agentic framework that turns user-level ODE and PDE descriptions into solver-backed neural-operator pipelines. The workflow links problem specification, data generation, operator training, and checkpoint-based inference. A stateful input graph converts multi-turn natural-language input and user edits into validated problem specifications. The data-generation module then samples parameters, solves the configured governing-equation with FEniCSx finite-element back
arXiv 25d ago Research Agents & autonomy

Look-Ahead-Freedom as Temporal Non-Interference: A Verifiable Correctness Property for Backtesting and Agentic Trading Pipelines

Look-ahead bias (using information from after a decision epoch to make the decision at that epoch) is the dominant way a backtest or a machine-learning evaluation flatters a system that will disappoint in deployment. The field manages it with construct-specific recipes and empirical detectors, which are sound only channel by channel and certify nothing by their silence. We show that look-ahead-freedom is a formal property in disguise: fixing an epoch, the demand that the future not influence the
arXiv cs.CR (AI security) 25d ago Research Bias & fairnessAgents & autonomy

The Robots Are Here

Unitree's advantage
ChinaTalk 25d ago Field notes Agents & autonomy

DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-level planning interface in VLM-based VLAs for decomposing coarse instructions. Such decomposition becomes important when household tasks involve complex multi-step goals, where coarse user commands need to be converted i
arXiv 25d ago Research Agents & autonomy
← Newer Older →