Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
Reliable and Developer-Aligned Evaluation of Agents for Software Engineering
Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their accelerated adoption, existing evaluation techniques are limited due to their fragmented nature and distorted projection of true model capabilities, often obtained from hypothetical syntactic scenarios. This research aims to bridge this gap by providing a comprehensive e
China’s Answer to AI Sticker Shock
Corporate America is starting to balk at the cost of AI agents. A cheap alternative from China looks more tempting than ever.
From Dashboards to Agents: The Future of Data Visualization
Data visualizations can inform, explain, and sway public opinion and policy decisions. This course imparts design thinking and data ethics frameworks, along with practical software skills, to ...
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
Large language model coding agents increasingly perform open-ended data modeling and analysis. These agents are stochastic and adaptive, and therefore their autonomous model discovery behavior cannot be adequately characterized by a single benchmark run. In this work, we propose an experimental design and analysis framework for systematically evaluating this discovery process, quantifying its variability, and identifying important factors. The proposed framework treats these agents as stochastic
'GitLost' Flaw Leaks Private Data From GitHub's Agentic Workflows
The flaw allows an unauthenticated attacker to craft a GitHub Issue in an org's public repository and then silently pull data from its private repos, too.
AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters
Max single-threaded CPUs at scale are a new category of CPUs built for the agentic AI era. Across the creation and deployment of an agentic system, the CPU is on the critical path for reasoning, response time and learning. CPUs are the processor which executes the work the AI model commands: the tool calling, code […]
Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction
While personalisation is becoming a defining capability in human-robot interaction (HRI), the existing literature on responsible personalisation remains fragmented, offering isolated accounts of ethical risks without a structured understanding of how they emerge across interaction contexts. This gap is particularly critical in HRI, where robots' embodiment and social presence can amplify and reshape such risks or generate new types of risks. We present a lifecycle-based and context-sensitive fra
SPEAR: A Simulator for Photorealistic Embodied AI Research
Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by introducing SPEAR: A Simulator for Photorealistic Embodied AI Research. At its core, SPEAR is a Python library that can connect to, and programmatically control, any Unreal Engine (UE) application via a modular plugin architecture. SPEAR expo
Accelerating science and medicine with collaborative agents
Google DeepMind’s Vivek Natarajan on porting AlphaGo’s self-play recipe into science and medicine, via the AI co-scientist and AMIE. From RAAIS 2026.
From Coding Robots to Speed Networking on a UFO: Day One at AI for Good’s Summit
GENEVA, 7 July 2026 — The Youth Zone at the AI for Good Global Summit 2026 commences its program today with a robotics competition, a series of hands-on artificial intelligence seminars, and a policy discussion on the skills that classrooms should prioritise as AI becomes more ingrained in daily life. A continuum approach to AI literacy, rather than a single fixed curriculum, is reflected in the participation of children as young as six and young adults. The post From Coding Robots to Speed Netw
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement. In this paper, we investigate deliberative large language model (LLM) agents under partially observable joint decision-making tasks. We formalize deliberative collaboration as a cooperative joint decision problem with partial and asymmetric observations, and introduce a scalable benchmark that instantiates this problem across multiple
Glass crashes slashed? Ant Group embodied AI unit claims breakthrough in robot sensing
Robbyant, the embodied artificial intelligence arm of Chinese fintech giant Ant Group, launched a new vision model that it claims can help robots overcome a long-standing challenge: accurately perceiving glass, mirrors and transparent objects. The unit of Hangzhou-based Ant Group on Tuesday unveiled its next-generation spatial perception model, LingBot-Depth 2.0, alongside a new foundational visual model called LingBot-Vision, as AI labs race to equip machines with the “brains” required to...
The foundational elements of AI architecture that IT leaders need to scale
With the rapid progress of AI capabilities and the move to agentic systems, organizations are expanding their use cases as the technology continues to grow. That constant evolution also introduces risk, leaving IT leaders to wonder which investments will prove valuable even six months into the future. Returning to the foundational elements of AI architecture—the…
How AI could enable autonomous robot workers in workplaces—and maybe homes
Top robotics researchers and founders explain how robot autonomy is evolving.
From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations
Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face practical limits on control and replication. Meanwhile, LLM-based social simulations are typically behavior-driven and lack theory-aligned environments for modeling Putnam's core propositions. To address these gaps, we introduce SocaSim, an LLM-based multi-agent simulation framework to study Putnam's Social Capital Theory from theoretical blueprin
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are ill-equipped to support given its evolving, interactive, and context-dependent nature. In this paper, we introduce Prompt Coach (PC), an agentic tutor that helps developers learn how to craft high-quality code-generation prompts through Socratic guidance embedded in-flow within their IDE. PC evaluates prompt quality across multiple dimensions and surfaces targe
Huawei’s new computing cluster, world’s first AI agent phone to debut at China AI summit
The coming World Artificial Intelligence Conference (WAIC) is set to feature major new product releases, such as Huawei Technologies’ next-generation computing cluster, as China doubles down on AI in the global technology race. This year’s WAIC, the ninth since 2018 and running from July 17 to 20 in Shanghai, would include the first physical display of Huawei’s Atlas 950 SuperPoD, Tang Wenkan, director of the Shanghai Municipal Commission of Economy and Informatisation, said at a press briefing.
Intelligence is Free, Now What? Data Systems for, of, and by Agents
... government of the people, by the people, for the people ... — Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly. GPT-4-class capabilities cost roughly $30 per million tokens in early 2023; today the same runs under $1 , and some providers are pushing costs below $0.10 . Across benchmarks, inference prices have fallen between 9x and 900x per year , with a median decline near 50x. Even frontier models are getting dramatically cheaper ea
Expanding Managed Agents in Gemini API: background tasks, remote MCP and more
We’re announcing new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.
Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems
AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost -- for instance, marketing agents may post misleading content as a result of competing for engagement on social media. Human societies address such problems through norms that constrain acceptable behavior, supported by enforcement mechanisms that detect and penalize violations. Motiva
NVIDIA and Hugging Face Bring New Models and Frameworks to LeRobot for the Open Robotics Community
Open source AI has shown how quickly developers can innovate when models, data and tools are shared. Robotics has the same opportunity, but advancements in physical AI development can still be gated by costly and fragmented resources, from large datasets and robot foundation models to simulation, compute and validation tools. NVIDIA and Hugging Face are […]
Security and Privacy in Agentic AI: Grand Challenges and Future Directions
We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought together thirty leading international experts from academia, industry, and government to engage in focused discussions and collaborative exercises on the emerging risks associated with the growing agency of AI.
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of
China records most new unicorn start-ups in 5 years as AI and robotics boom
China’s innovation ecosystem has witnessed a resurgence, minting 67 new unicorn start-ups in the first half of 2026 – the biggest increase in almost five years – as AI and robotics kick off a new investment cycle. The growth translates into an average of one new unicorn – private companies valued at US$1 billion or more – in less than every three days and was the highest since the second half of 2021 when 76 new unicorns were created, according to a Monday report by ITJuzi, a start-up...
Social cognitive architecture for NPC groups: integration of transformer theory of mind and hierarchical reinforcement learning
Non-Player Characters (NPCs) require social cognition to enable intelligent and interactive behaviours within virtual environments. In gaming and other multi-agent systems, current NPC models often fall short in social intelligence and coordination because they cannot infer or anticipate the mental states of other agents. To address this gap, this paper introduces a novel social cognitive architecture that integrates Hierarchical Reinforcement Learning (HRL) with a Transformer-based Theory of Mi
Giving Meaning to Technological Artifacts in One’s Life: An Aspect of Personal Autonomy in a Technological Society
In today’s world, where advanced technology such as artificial intelligence (AI) is reshaping human society, how can individuals’ autonomy in the use of such technology be understood? This paper focuses on the meanings of technological artifacts in individuals’ lives—that is, the subjective interpretation of their function within one’s plan for use—and attempts to formalize a set of human capacities to actively give meaning to technological artifacts as a form of personal autonomy. This study sh
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for RL training, thus failing to capture web diversity. We propose Weblica (Web Replica), a framework for constructing reproducible and scalable web environments. Our framework leverages 1) HTTP-level caching to capture and replay stabl
Robots available for rent: But what can they do?
Robotics tech is changing fast, so for many it makes sense to rent a robot.
Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies
Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurement. While cryptographic techniques protect explicitly disclosed constraint values, they fail to address a subtler threat: behavioral privacy leakage, where an adversary infers private constraints from observable negotiation dynamics such as concession trajectories, timing, and convergence patterns. This paper investigates behavioral differential privacy in multi-round negotiation protoc
When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents
Personal AI agents powered by large language models can reason and act using available tools to access emails, manage calendars, and push code to remote repositories, all with minimal oversight. When augmented with long-term memory, an agent can recall specific details relevant to the current task, reducing the need for large context windows. Currently, long-term memory agents tend to fall into two distinct domains: conversational and action-planning agents. Personal assistant agents sit at the
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations. Hierarchical dual-system methods address this but suffer from a gap between high-level planning semantics and low-level execution kinematics. We introduce Cortex, a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negotiate with services. Existing benchmarks evaluate tool use, web navigation, desktop control, personalization, recommendation, and evolving context, but rarely ask whether an agent preserves user sovereignty: advancing the user's current interests while respecting privacy, consent, evidence, user burden, and resistance to manipulative incentives. W
Multiplayer Interactive World Models with Representation Autoencoders
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, t
BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking
Public battery aging datasets are a critical asset for advanced health management, but their practical use is often limited by inconsistent formats, unclear schemas, and metadata scattered across repositories and publications. Current curation remains largely manual and hard to reproduce, while general-purpose data integration tools miss the domain-specific semantics of electrochemical time-series data. We present BatteryLake, a governed data lakehouse that turns raw public battery data into ben
JadePuffer: The First Complete LLM-Driven Ransomware Attack
An "agentic threat actor" successfully exploited a Langflow flaw to steal data from a production database server and encrypt other systems.
Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets
In liberalised railway systems, operators must set prices dynamically in an environment with partial observability, as they retain private information about their objectives and performance, where regulatory constraints prohibit communication or direct information exchange between competitors to prevent explicit collusion. Consequently, agents must learn to infer strategic interactions only from observable market data which presents a significant challenge for multi-agent reinforcement learning,
PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference
We present PDEFlow, an autonomous agentic framework that turns user-level ODE and PDE descriptions into solver-backed neural-operator pipelines. The workflow links problem specification, data generation, operator training, and checkpoint-based inference. A stateful input graph converts multi-turn natural-language input and user edits into validated problem specifications. The data-generation module then samples parameters, solves the configured governing-equation with FEniCSx finite-element back
Look-Ahead-Freedom as Temporal Non-Interference: A Verifiable Correctness Property for Backtesting and Agentic Trading Pipelines
Look-ahead bias (using information from after a decision epoch to make the decision at that epoch) is the dominant way a backtest or a machine-learning evaluation flatters a system that will disappoint in deployment. The field manages it with construct-specific recipes and empirical detectors, which are sound only channel by channel and certify nothing by their silence. We show that look-ahead-freedom is a formal property in disguise: fixing an epoch, the demand that the future not influence the
The Robots Are Here
Unitree's advantage
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-level planning interface in VLM-based VLAs for decomposing coarse instructions. Such decomposition becomes important when household tasks involve complex multi-step goals, where coarse user commands need to be converted i