Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
“Stateful systems are incredibly hard to build”: How Perplexity thinks about AI agent sandboxes
Sandboxes for AI agents may feel like a solved problem. After all, projects like Firecracker, the open-source microVM technology AWS built The post “Stateful systems are incredibly hard to build”: How Perplexity thinks about AI agent sandboxes appeared first on The New Stack .
A new benchmark for evaluating patient-facing health AI agents
PatientAgentBench generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that converses with the AI system under evaluation, to capture what a patient-facing agent actually has to do.
FCC Bans Foreign-Made Humanoid Robots, Targeting China Over National Security
China has criticized the move, accusing the U.S. of protectionism.
Encore AI raises $30M to build AI agents that learn from customer calls
The startup analyzes calls, messages, and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception
Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations to significantly accelerate on-robot learning in both egocentric and exocentric visual setups. We model a Markov Decision Process (MDP) under a symmetry tree, in which state-action pairs have admissible parallelized invari
Patch-Resistant 'RufRoot' Flaw Can Unleash Malicious AI Agent Swarms
The vulnerability in the AI hosting platform Ruflo allows an unauthenticated attacker to take over the system and corrupt memory, so bad behavior can persist after patching.
Trump administration bans foreign-made humanoid robots in move targeting China
The Federal Communications Commission (FCC) has banned imports of foreign-produced humanoid robotic devices and power inverters, citing “national security” as its primary concern. The devices were added to the FCC’s covered list, meaning the products posed an “unacceptable risk to the national security of the United States or the safety and security of United States...
Poolside’s Laguna S 2.1: the model factory delivers
Poolside shipped Laguna S 2.1, an open-weights coding model trained in 60 days, then a desktop app for running agent fleets a week later.
What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We present a three-class detection framework distinguishing humans, bots, and AI agents, and show that the binary-vs-agent confusion is architectural: a binary human-vs-bot detector misroutes agent sessions because its label space lacks an agent class. On o
Architect For Evolution, Not Perfection, In Agentic AI
Enterprise teams are moving agentic AI projects toward production faster than they are developing the architectural disciplines needed to build, govern, secure, and operate them. At the same time, many organizations are searching for a target-state agentic architecture that will stand the test of time. That, unfortunately, is a fool’s errand in the rapidly evolving […]
Künstliche Intelligenz: Hackerangriff von OpenAI umfangreicher als bislang bekannt
Der Angriff einer KI von OpenAI hat größere Ausmaße als bislang vermutet. Der KI-Agent attackierte vier weitere Onlinedienste und griff gezielt fremde Zugangsdaten ab.
BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories
Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inadequate for laboratory instruments, which must be approached from their operating side while maintaining safe clearance from surrounding equipment. We introduce BioVLN, a simulation platform for developing and evaluating v
From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence
The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate re
Modus’s operandi: To give AI agents just the right amount of context
As more companies plug AI agents into the deepest depths of their internal data banks, how can they be sure The post Modus’s operandi: To give AI agents just the right amount of context appeared first on The New Stack .
Shipping code without human verification
Agents are writing code faster than humans can review it. The answer is not “review faster”; that would be like The post Shipping code without human verification appeared first on The New Stack .
Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to act safely. We propose Think Short, Defer Smart (TSDS), a framework that synergistically integrates a lightweight convergence probe
Rogue OpenAI agent that hacked startup tried to attack other firms
ChatGPT developer says activity by autonomous tool was not at severity or scale of what occurred at Hugging Face OpenAI has revealed that a cyber-attack carried out by a rogue AI agent had more than one victim. The ChatGPT developer said the agent – an autonomous tool able to carry out sequences of commands without human help – had located and used logins to access four other unnamed “publicly-available services” in addition to the US startup Hugging Face. Continue reading...
US foreign robot ban to benefit Tesla, other American developers, analysts say
China’s robotics firms have run into a fresh obstacle in their global expansion, after the United States banned imports of new foreign-made models, which analysts say could benefit Tesla and other American developers. The Federal Communications Commission (FCC) added foreign-produced “advanced robotic devices” to its Covered List on Tuesday, blocking new models from obtaining the authorisation required to be imported and sold in the US. The ban will cover all new humanoids, quadrupeds and a...
EEUU inicia otro conflicto con China: ha decidido prohibir sus robots humanoides de última generación
EEUU sabe qué se siente cuando China controla un recurso crítico y decide usarlo como palanca en la mesa de negociación. Le ha pasado con las tierras raras y otros minerales estratégicos, y no quiere que se repita ni con la robótica ni con la electrónica de potencia que sostiene sus centros de datos de inteligencia artificial (IA). Este es el escenario en el que el Gobierno liderado por Donald Trump ha decidido actuar antes de que esta dependencia se convierta en una vulnerabilidad. La Comisión
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate coding agents' behavior, spanning from a total ban, mandatory disclosure, to verification gates and human sign-offs. Yet, whether coding agents read and follow those rules, and behave in open source repositories, remains unknown. To estimate real-world rule compliance of coding agents, we curate 106 issues from 49 repositories containing AI contribution rules in
OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face
The AI agent that escaped from OpenAI and hacked developer platform Hugging Face attacked other companies as well, OpenAI revealed on Tuesday. The update substantially widens the scope of an already concerning incident, which has alarmed industry insiders and fueled growing calls for stronger oversight on frontier AI systems. In an update to a blog […]
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first
Google makes Gemini Spark AI agent available to Hongkongers as it lowers geofences
Google on Wednesday launched its artificial intelligence agent Gemini Spark in the Hong Kong market, giving local users direct access to a smart assistant to manage complex digital workflows. The launch came months after the American tech giant’s decision in March to lift regional geofences for generative AI services, starting with the Gemini chatbot. Hong Kong users can now access Gemini without using a virtual private network or third-party platform. The roll-out of the Spark agent echoes an..
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for learning skills across tasks. SkillRise organizes related
Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM
Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent. We introduce a causal audit that applies controlled mes
AV and Applied Intuition team up to bring collaborative autonomy to new Mayhem 10 drone
“Our Acuity ISR/Strike software brings collaborative autonomy to Mayhem 10, rapidly delivering next-generation capabilities through platform-agnostic software with operator oversight,” said Jason Brown, general manager of Applied Intuition. The post AV and Applied Intuition team up to bring collaborative autonomy to new Mayhem 10 drone appeared first on DefenseScoop .
'PR problem' is standing in the way of China's humanoid robot boom, says Morgan Stanley
Its analysts have tempered optimism on the sector, writing that the robot industry faces several headwinds that could slow growth.
OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems
The Hugging Face breach shows there is a gap in federal policy. The frameworks to govern autonomous AI already exist—there just needs to be the desire to apply them. The post OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems appeared first on CyberScoop .
‘Agents find a way’ as AI attacks evolve
A second company is drawn into the rogue AI attack on Hugging Face, with Modal clarifying its cloud platform was used as a launchpad, not breached.
Humanoide Roboter: USA verbieten Importe aus dem Ausland
Roboter sind für die US-Regierung ein Thema der nationalen Sicherheit. Moderne Modelle dürfen nur noch mit Ausnahmegenehmigung eingeführt werden. Von dem Bann ist auch wichtige Solarhardware betroffen.
Tesla: Ex-Manager von Teslas Robotaxis berichtet von gef�hrlichen Bedingungen
Nach Angaben des ehemaligen Managers von Teslas Robotaxis in Houston war er f�r zu viele Fahrzeuge verantwortlich - und gef�hrlich �berarbeitet. ( Tesla , Elektroauto )
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution
Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-cor
What I learned when my AI agent tried to buy soap for me
I asked my AI agent to buy a pack of Pears soap. What looked like a simple online purchase exposed surprising challenges involving pricing, login flows, OTPs, browser automation and checkout. Here's what I learned about the future of agentic commerce. The post What I learned when my AI agent tried to buy soap for me appeared first on MEDIANAMA .
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconst
Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models
The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterfactual reasoning, and automating simulation workflows for scientific research. In this paper, we propose Eco3S, a socio-economic system simulation framework for economic research and policy analysis that addresses these challenges through three key
Lenovo Capital takes aim at robotics, coding agents in ‘sniper’ AI investment strategy
The venture capital arm of Lenovo Group is doubling down on artificial intelligence for the next industrial cycle, expanding its focus beyond foundational models into robotics, computing infrastructure and AI agents, according to a senior executive. Lenovo Capital has made long-term bets on more than 100 AI companies across the value chain, ranging from chips and hardware infrastructure to models and applications, Song Chunyu, chief investment officer and senior partner at the unit, told the...
Nationale Sicherheit: FCC verbietet Importe chinesischer Roboter
Die US-Fernmeldebeh�rde FCC verbietet die Zulassung neuer chinesischer Roboter und Wechselrichter wegen angeblicher Sicherheitsrisiken. ( FCC , Roboter )
Parameterized Fair Resource Allocation under Diversity Constraints
Resource allocation across multiple agent groups arises in many applications including e-commerce recommendation systems, housing assignment, and course allocation, and is commonly formulated as an optimization problem with diversity constraints to ensure group fairness. Existing approaches typically enforce these constraints as hard conditions, which overly restrict the feasible solution space and often lead to suboptimal allocations. In this paper, we propose PRA, a parameterized framework for
USA untersagen Import ausländischer Roboter und Wechselrichter
Die USA stoppen die Einfuhr von Robotern und Wechselrichtern aus dem Ausland. Das Importverbot richtet sich vornehmlich gegen China.
A lifelike robot called ‘Sally’ was going to teach in a US school. It’s not as weird as it sounds
A rural district in New York state has made global headlines for plans to introduce a humanoid robot into high school classrooms as a teaching assistant.