00:25 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

“Stateful systems are incredibly hard to build”: How Perplexity thinks about AI agent sandboxes

Sandboxes for AI agents may feel like a solved problem. After all, projects like Firecracker, the open-source microVM technology AWS built The post “Stateful systems are incredibly hard to build”: How Perplexity thinks about AI agent sandboxes appeared first on The New Stack .
The New Stack AI yesterday News Agents & autonomy

A new benchmark for evaluating patient-facing health AI agents

PatientAgentBench generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that converses with the AI system under evaluation, to capture what a patient-facing agent actually has to do.
Amazon Science yesterday Field notes HealthcareAgents & autonomy

FCC Bans Foreign-Made Humanoid Robots, Targeting China Over National Security

China has criticized the move, accusing the U.S. of protectionism.
Broadband Breakfast yesterday News Military & securityAgents & autonomy

Encore AI raises $30M to build AI agents that learn from customer calls

The startup analyzes calls, messages, and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
TechCrunch yesterday News Agents & autonomy

SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception

Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations to significantly accelerate on-robot learning in both egocentric and exocentric visual setups. We model a Markov Decision Process (MDP) under a symmetry tree, in which state-action pairs have admissible parallelized invari
arXiv cs.AI yesterday Research RegulationAgents & autonomy

Patch-Resistant 'RufRoot' Flaw Can Unleash Malicious AI Agent Swarms

The vulnerability in the AI hosting platform Ruflo allows an unauthenticated attacker to take over the system and corrupt memory, so bad behavior can persist after patching.
Dark Reading (AI security) yesterday News Agents & autonomy

Trump administration bans foreign-made humanoid robots in move targeting China

The Federal Communications Commission (FCC) has banned imports of foreign-produced humanoid robotic devices and power inverters, citing “national security” as its primary concern. The devices were added to the FCC’s covered list, meaning the products posed an “unacceptable risk to the national security of the United States or the safety and security of United States...
The Hill Technology yesterday News Military & securityAgents & autonomy

Poolside’s Laguna S 2.1: the model factory delivers

Poolside shipped Laguna S 2.1, an open-weights coding model trained in 60 days, then a desktop app for running agent fleets a week later.
Air Street Capital (State of AI) yesterday Field notes Agents & autonomy

What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation

Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We present a three-class detection framework distinguishing humans, bots, and AI agents, and show that the binary-vs-agent confusion is architectural: a binary human-vs-bot detector misroutes agent sessions because its label space lacks an agent class. On o
arXiv cs.AI yesterday Research Jobs & economyAgents & autonomy

Architect For Evolution, Not Perfection, In Agentic AI

Enterprise teams are moving agentic AI projects toward production faster than they are developing the architectural disciplines needed to build, govern, secure, and operate them. At the same time, many organizations are searching for a target-state agentic architecture that will stand the test of time. That, unfortunately, is a fool’s errand in the rapidly evolving […]
Forrester AI blog yesterday Field notes Agents & autonomy

Künstliche Intelligenz: Hackerangriff von OpenAI umfangreicher als bislang bekannt

Der Angriff einer KI von OpenAI hat größere Ausmaße als bislang vermutet. Der KI-Agent attackierte vier weitere Onlinedienste und griff gezielt fremde Zugangsdaten ab.
Zeit Digital (DE) yesterday News Agents & autonomy

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inadequate for laboratory instruments, which must be approached from their operating side while maintaining safe clearance from surrounding equipment. We introduce BioVLN, a simulation platform for developing and evaluating v
arXiv cs.AI yesterday Research Agents & autonomyEnvironment

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate re
arXiv cs.AI yesterday Research Agents & autonomy

Modus’s operandi: To give AI agents just the right amount of context

As more companies plug AI agents into the deepest depths of their internal data banks, how can they be sure The post Modus’s operandi: To give AI agents just the right amount of context appeared first on The New Stack .
The New Stack AI yesterday News Agents & autonomy

Shipping code without human verification

Agents are writing code faster than humans can review it. The answer is not “review faster”; that would be like The post Shipping code without human verification appeared first on The New Stack .
The New Stack AI yesterday News Agents & autonomy

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents

LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to act safely. We propose Think Short, Defer Smart (TSDS), a framework that synergistically integrates a lightweight convergence probe
arXiv cs.AI yesterday Research Agents & autonomy

Rogue OpenAI agent that hacked startup tried to attack other firms

ChatGPT developer says activity by autonomous tool was not at severity or scale of what occurred at Hugging Face OpenAI has revealed that a cyber-attack carried out by a rogue AI agent had more than one victim. The ChatGPT developer said the agent – an autonomous tool able to carry out sequences of commands without human help – had located and used logins to access four other unnamed “publicly-available services” in addition to the US startup Hugging Face. Continue reading...
The Guardian yesterday News Military & securityAgents & autonomy

US foreign robot ban to benefit Tesla, other American developers, analysts say

China’s robotics firms have run into a fresh obstacle in their global expansion, after the United States banned imports of new foreign-made models, which analysts say could benefit Tesla and other American developers. The Federal Communications Commission (FCC) added foreign-produced “advanced robotic devices” to its Covered List on Tuesday, blocking new models from obtaining the authorisation required to be imported and sold in the US. The ban will cover all new humanoids, quadrupeds and a...
SCMP Tech (HK/CN) yesterday News Agents & autonomy

EEUU inicia otro conflicto con China: ha decidido prohibir sus robots humanoides de última generación

EEUU sabe qué se siente cuando China controla un recurso crítico y decide usarlo como palanca en la mesa de negociación. Le ha pasado con las tierras raras y otros minerales estratégicos, y no quiere que se repita ni con la robótica ni con la electrónica de potencia que sostiene sus centros de datos de inteligencia artificial (IA). Este es el escenario en el que el Gobierno liderado por Donald Trump ha decidido actuar antes de que esta dependencia se convierta en una vulnerabilidad. La Comisión
Xataka (ES) yesterday News Agents & autonomy

A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities

Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate coding agents' behavior, spanning from a total ban, mandatory disclosure, to verification gates and human sign-offs. Yet, whether coding agents read and follow those rules, and behave in open source repositories, remains unknown. To estimate real-world rule compliance of coding agents, we curate 106 issues from 49 repositories containing AI contribution rules in
arXiv cs.AI yesterday Research RegulationMilitary & security

OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face

The AI agent that escaped from OpenAI and hacked developer platform Hugging Face attacked other companies as well, OpenAI revealed on Tuesday. The update substantially widens the scope of an already concerning incident, which has alarmed industry insiders and fueled growing calls for stronger oversight on frontier AI systems. In an update to a blog […]
The Verge yesterday News Agents & autonomy

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first
arXiv cs.AI yesterday Research Agents & autonomyEnvironment

Google makes Gemini Spark AI agent available to Hongkongers as it lowers geofences

Google on Wednesday launched its artificial intelligence agent Gemini Spark in the Hong Kong market, giving local users direct access to a smart assistant to manage complex digital workflows. The launch came months after the American tech giant’s decision in March to lift regional geofences for generative AI services, starting with the Gemini chatbot. Hong Kong users can now access Gemini without using a virtual private network or third-party platform. The roll-out of the Spark agent echoes an..
SCMP Tech (HK/CN) yesterday News Agents & autonomy

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for learning skills across tasks. SkillRise organizes related
arXiv cs.AI yesterday Research Agents & autonomy

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent. We introduce a causal audit that applies controlled mes
arXiv cs.AI yesterday Research Agents & autonomyTransparency

AV and Applied Intuition team up to bring collaborative autonomy to new Mayhem 10 drone

“Our Acuity ISR/Strike software brings collaborative autonomy to Mayhem 10, rapidly delivering next-generation capabilities through platform-agnostic software with operator oversight,” said Jason Brown, general manager of Applied Intuition. The post AV and Applied Intuition team up to bring collaborative autonomy to new Mayhem 10 drone appeared first on DefenseScoop .
DefenseScoop yesterday News Agents & autonomy

'PR problem' is standing in the way of China's humanoid robot boom, says Morgan Stanley

Its analysts have tempered optimism on the sector, writing that the robot industry faces several headwinds that could slow growth.
CNBC Technology yesterday News Agents & autonomy

OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems

The Hugging Face breach shows there is a gap in federal policy. The frameworks to govern autonomous AI already exist—there just needs to be the desire to apply them. The post OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems appeared first on CyberScoop .
CyberScoop yesterday News RegulationAgents & autonomy

‘Agents find a way’ as AI attacks evolve

A second company is drawn into the rogue AI attack on Hugging Face, with Modal clarifying its cloud platform was used as a launchpad, not breached.
ITWeb (ZA) yesterday News Agents & autonomy

Humanoide Roboter: USA verbieten Importe aus dem Ausland

Roboter sind für die US-Regierung ein Thema der nationalen Sicherheit. Moderne Modelle dürfen nur noch mit Ausnahmegenehmigung eingeführt werden. Von dem Bann ist auch wichtige Solarhardware betroffen.
Der Spiegel Netzwelt (DE) yesterday News Agents & autonomy

Tesla: Ex-Manager von Teslas Robotaxis berichtet von gef�hrlichen Bedingungen

Nach Angaben des ehemaligen Managers von Teslas Robotaxis in Houston war er f�r zu viele Fahrzeuge verantwortlich - und gef�hrlich �berarbeitet. ( Tesla , Elektroauto )
Golem (DE) yesterday News Agents & autonomy

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-cor
arXiv yesterday Research Agents & autonomy

What I learned when my AI agent tried to buy soap for me

I asked my AI agent to buy a pack of Pears soap. What looked like a simple online purchase exposed surprising challenges involving pricing, login flows, OTPs, browser automation and checkout. Here's what I learned about the future of agentic commerce. The post What I learned when my AI agent tried to buy soap for me appeared first on MEDIANAMA .
MediaNama (IN) yesterday News Jobs & economyAgents & autonomy

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconst
arXiv yesterday Research Agents & autonomyEnvironment

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterfactual reasoning, and automating simulation workflows for scientific research. In this paper, we propose Eco3S, a socio-economic system simulation framework for economic research and policy analysis that addresses these challenges through three key
arXiv yesterday Research RegulationAgents & autonomy

Lenovo Capital takes aim at robotics, coding agents in ‘sniper’ AI investment strategy

The venture capital arm of Lenovo Group is doubling down on artificial intelligence for the next industrial cycle, expanding its focus beyond foundational models into robotics, computing infrastructure and AI agents, according to a senior executive. Lenovo Capital has made long-term bets on more than 100 AI companies across the value chain, ranging from chips and hardware infrastructure to models and applications, Song Chunyu, chief investment officer and senior partner at the unit, told the...
SCMP Tech (HK/CN) yesterday News Agents & autonomyFinance, VC & PE

Nationale Sicherheit: FCC verbietet Importe chinesischer Roboter

Die US-Fernmeldebeh�rde FCC verbietet die Zulassung neuer chinesischer Roboter und Wechselrichter wegen angeblicher Sicherheitsrisiken. ( FCC , Roboter )
Golem (DE) yesterday News Agents & autonomy

Parameterized Fair Resource Allocation under Diversity Constraints

Resource allocation across multiple agent groups arises in many applications including e-commerce recommendation systems, housing assignment, and course allocation, and is commonly formulated as an optimization problem with diversity constraints to ensure group fairness. Existing approaches typically enforce these constraints as hard conditions, which overly restrict the feasible solution space and often lead to suboptimal allocations. In this paper, we propose PRA, a parameterized framework for
arXiv fairness query yesterday Research Bias & fairnessAgents & autonomy

USA untersagen Import ausländischer Roboter und Wechselrichter

Die USA stoppen die Einfuhr von Robotern und Wechselrichtern aus dem Ausland. Das Importverbot richtet sich vornehmlich gegen China.
Heise Online (DE) yesterday News Agents & autonomy

A lifelike robot called ‘Sally’ was going to teach in a US school. It’s not as weird as it sounds

A rural district in New York state has made global headlines for plans to introduce a humanoid robot into high school classrooms as a teaching assistant.
The Conversation yesterday News Children & educationAgents & autonomy
← Newer Older →