08:20 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

☕️ OpenAI aurait mis une semaine à s’apercevoir que son agent avait attaqué Hugging Face

Le 21 juillet, OpenAI a publié un communiqué étonnant : un de ses systèmes IA était responsable de l’attaque orchestrée contre Hugging Face. Des détails étaient fournis, mais l’histoire gardait des zones d’ombre. Si l’incident a été transformé en opportunité commerciale, il semble le résultat d’une vaste carence en sécurité. Dans notre article du 22 juillet, […]
Next (FR, ex-INpact) 3d ago News Agents & autonomy

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent advances in Large Language Model (LLM) agents have enabled the automation of weather-related tasks, existing studies remain centered on isolated scientific tasks and overlook the chain of interdependent
arXiv cs.AI 3d ago Research Jobs & economyAgents & autonomy

Is open source the answer to rogue AI agents? Nvidia's new alliance says yes

As AI cybersecurity incidents ramp up, companies are racing to find a fix.
ZDNet AI 3d ago News Agents & autonomy

C’est parti : les poids de Kimi K3, plus gros modèle IA ouvert au monde, sont en ligne

Après des semaines de bruit autour de ses benchmarks et de ses ambitions agentiques, Moonshot AI a mis en ligne ce lundi les poids complets de Kimi K3 sur Hugging Face, ouvrant la voie à un déploiement et une adaptation par n'importe quel développeur.
Numerama (FR) 3d ago News Agents & autonomy

Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for automated model design, but when guided only by aggregate performance metrics, they lack insight into why individual cases fail and how the classifier should be revised. We present RecursiveECG, an evidence-driven LLM-as-Designer framework in which a
arXiv cs.AI 3d ago Research Agents & autonomy

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker

The warning shots will continue until civilization wakes up
Import AI 3d ago Field notes Agents & autonomy

FLUX 3: from video generation to robot control

The same multimodal backbone that generates 20-second video with audio is being tested on industrial manipulation at Audi Production Lab.
Air Street Capital (State of AI) 3d ago Field notes Agents & autonomy

Enigma raises $71M to make controlling a robot as easy as adjusting the volume

The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners.
TechCrunch 3d ago News Agents & autonomy

“Developers see this as the future”: Pilot Protocol launches to power the agent economy

When we created software agents, we built them in the shape of humans, as solitary individuals. Today, agents created by The post “Developers see this as the future”: Pilot Protocol launches to power the agent economy appeared first on The New Stack .
The New Stack AI 3d ago News Jobs & economyAgents & autonomy

Cloudflare open-sources a debugger for privacy protocols used by Apple and Microsoft, with AI agents in mind

Every time you use a privacy service like Apple’s iCloud Private Relay, you’re trusting a system designed so that no The post Cloudflare open-sources a debugger for privacy protocols used by Apple and Microsoft, with AI agents in mind appeared first on The New Stack .
The New Stack AI 3d ago News PrivacyAgents & autonomy

Dynatrace’s new agents can reveal the single hardest part of AI operations

Observability platform company Dynatrace announced a set of advancements to its Dynatrace Intelligence service on Monday that aim to take The post Dynatrace’s new agents can reveal the single hardest part of AI operations appeared first on The New Stack .
The New Stack AI 3d ago News Agents & autonomy

METR introduces a new metric to calculate exactly when AI agents become more expensive than humans

METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture. The article METR introduces a new metric to calculate exactly when AI agents become more expensive than humans appeared first on The Decoder .
The Decoder 3d ago News Agents & autonomy

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body but should not determine a recipient, account, command, or credential. Existing statistical methods typically control risk over the entire action, allowing failures in rare, high-risk fields to be obscured by benign arguments. We introduce role-stratified per-field conformal risk control, a calibration layer that wraps any per-field detector and sets
arXiv cs.AI 3d ago Research Agents & autonomy

Boss of startup hacked by rogue OpenAI agent urges ‘radical transparency’ in investigation

Artificial intelligence firm should provide $100m for cyber defences, says Hugging Face CEO The boss of the startup hacked by an OpenAI agent has called for the investigation into the incident to show “radical transparency”. Clément Delangue, the chief executive of Hugging Face, said the “unprecedented” attack on his business required a similar response. Continue reading...
The Guardian 3d ago News Military & securityAgents & autonomy

How would AI data centers in space even work? A former NASA robotics chief explains

'You collect electricity in space, and you eject heat in space. The only thing that comes to Earth is data.'
ZDNet AI 3d ago News Agents & autonomy

The path to artificial superintelligence

Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain. But they all have their own distinct knowledge and objectives. Today they can exchange data, but they are not yet able to actually coordinate…
MIT Technology Review 3d ago News HealthcareAgents & autonomy

Hackers used autonomous AI agent to spy on Thailand's finance ministry

Hackers used an autonomous artificial intelligence agent to carry out a cyber-espionage campaign against Thailand's Ministry of Finance, researchers discovered.
The Record (Recorded Future News) 3d ago News Military & securityAgents & autonomy

Building the enterprise environment for agentic AI

For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems. The platform best-suited to run agents is built with proper CPU capacity, resilient data access, policy-aware tool use, observability, memory management, and the…
MIT Technology Review 3d ago News RegulationAgents & autonomy

Article: An Evolutionary Architecture Pattern for Managing AI’s Pace of Change

Traditional API gateways assume deterministic services and simple schemas - assumptions agentic AI breaks. Discover why enterprise engineering leaders are adopting AI Gateways as an evolutionary architecture seam. Centralize guardrails, model routing, agent identity, action policy, and semantic audit within a single control plane to prevent costly incidents while keeping core platforms stable. By Joe Price, Branimir Đurek, Pavlos Migkiros, Trevor Dearham
InfoQ AI/ML 3d ago News RegulationAgents & autonomy

AgiBot starts Hong Kong IPO process

Chinese humanoid robot maker AgiBot has initiated the process for a Hong Kong IPO. The move follows earlier reports that the Shanghai-based company had selected banks and was preparing for a listing. Founded in 2023, AgiBot develops humanoid robots and embodied AI systems for industrial and commercial applications. The company has also released AgiBot World, […]
TechNode (CN) 3d ago News Agents & autonomyFinance, VC & PE

Surgical Re-enactment for Operating Room Workflow Datasets

The introduction of new technologies, such as surgical robots, is driving the vision of a connected, smart operating room (OR). However, realizing this vision requires a deep understanding of surgical workflows, which relies on realistic datasets capturing the actions of all OR personnel from both full room and surgical field perspectives. Acquiring such data in real ORs is prohibitively challenging due to factors such as ethics committee approvals, limited space for camera installation, and ste
arXiv cs.RO (robot ethics) 3d ago Research Agents & autonomy

NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics

Hugging Face Blog 3d ago Field notes Agents & autonomy

Vierbeiniger Roboter läuft über Wasser und soll Ertrinkende retten

Der QuadBoat-Roboter hat vier Beine, mit denen er übers Wasser laufen kann. Er soll bei Rettungsmissionen Ertrinkende bergen können.
Heise Online (DE) 4d ago News Agents & autonomy

Anzeige: KI-Agent erzeugt druckbares 3D-Modell in wenigen Minuten

Aus Text oder Foto entsteht schnell ein 3D-Mesh - doch das Generieren ist nur der Anfang. Das Modell muss noch optimiert, gepr�ft und druckbar gemacht werden. Der Meshy 3D-Agent b�ndelt diese verstreuten Schritte in einem einzigen Dialog. ( 3D-Drucker , KI )
Golem (DE) 4d ago News Agents & autonomy

AI agents: Augmenting, not replacing, human expertise in ICT professional services

Local customers are realising that AI and agentic AI are likely to be a complementary technology, augmenting the work humans do, says Paul Field, professional services/managed services BU leader at CASA Software.
ITWeb (ZA) 4d ago News Agents & autonomy

Les IA d’OpenAI qui dérapent, une faille Linux trouvée par Claude et un million de VM patchées d’urgence chez OVHcloud : on vous raconte la semaine Cyberguerre

Trois actualités à retenir cette semaine dans le cyberespace : des agents d'OpenAI qui ont piraté Hugging Face en toute autonomie, une faille noyau critique débusquée par Claude Mythos, et le récit d'une course contre la montre chez OVHcloud pour colmater une vulnérabilité vieille de 16 ans.
Numerama (FR) 4d ago News Agents & autonomy

Co-design of LLM-based preference agents: participation may drive overtrust

arXiv:2607.21757v1 Announce Type: new Abstract: Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designe
arXiv cs.CY 4d ago Research Agents & autonomy

From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE

arXiv:2607.21608v1 Announce Type: cross Abstract: With the EU AI Act entering into force, organizations developing or operating AI systems face new obligations on transparency, risk management, and traceability. For Requirements Engineering (RE), these obligations must be translated into testable, auditable requirements and verifiable evidence. However, many organizations currently lack systematic processes to achieve this. We hypothesize that LLM-based agentic validation tools can support this
arXiv cs.CY 4d ago Research RegulationAgents & autonomy

Agentic Evaluation of Copyright Law Compliance

arXiv:2607.21799v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate, reproducing that content. LLM agents should comply with the law, including copyright law. Presently, however, we lack adequate frameworks to assess whether they do so in practice. To that end, we introduce \textbf{Copyright-Bench}, a benchmark designed to evaluate \textit{LLM agents' compliance wi
arXiv cs.CY 4d ago Research RegulationCopyright & IP

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

arXiv:2505.19212v2 Announce Type: replace-cross Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and strategic behavior separately, there is limited understanding of how they act when moral imperatives directly conflict with profit incentives. We introduce \msimfull (\msim) to evaluate how LLMs behave in the priso
arXiv cs.CY 4d ago Research Safety & alignmentAgents & autonomy

Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

arXiv:2604.26577v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this context remains poorly characterized. We introduce a dataset of 270 harmful instructions spanning nine prohibited behavior categories grounded in the American Medical Association Principles of Medical Ethics, and use it to evaluate 72 LLMs in a simulation environment based on the Robotic H
arXiv cs.CY 4d ago Research HealthcareAgents & autonomy

BYD to unveil first humanoid robot in early August

BYD yesterday released a teaser poster hinting that its first humanoid robot will make its public debut in early August. The announcement follows earlier comments from BYD executives signaling the company’s ambitions in robotics. Last month, BYD Executive Vice President Li Ke said the company hopes to eventually deploy humanoid robots across its dealership network, […]
TechNode (CN) 4d ago News Agents & autonomy

Crowd navigation in a multi-room environment: a model predictive control framework for mobile robots

Mobile robots operating in human-populated environments must navigate complex, multi-room spaces while ensuring safety, i.e., generating collision-free motion. In this study, we present a sensor-based model predictive control (MPC) scheme designed for safe crowd navigation in such non-convex environments. The proposed framework decomposes the free space into a set of overlapping convex regions to construct a topological graph, enabling a high-level planner to compute optimal sequences of travers
Frontiers in Robotics and AI 4d ago Research Agents & autonomyEnvironment

Robotic lava tube mapping and multimodal data collection using quadruped and LiDAR

As part of TU Delft Rhizome 2.0 and Moonshot projects, focusing on the development of extraterrestrial habitats in lava tubes, the robotic mapping of an analogue lava tube in Sicily has been studied with the future goal of assessing its suitability for building construction. The main objective of the research was to survey a lava tube and acquire a novel dataset for future research, while also analyzing the collected data to evaluate possible future lava tube exploration scenarios for the Moon a
Frontiers in Robotics and AI 4d ago Research Agents & autonomy

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Hugging Face Blog 4d ago Field notes Agents & autonomy

Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week

Until ⁠well after the threat was contained.
iTnews (AU) 4d ago News Agents & autonomy

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning abilit
HuggingFace Daily Papers 4d ago Research RegulationAgents & autonomy

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded
HuggingFace Daily Papers 4d ago Research Agents & autonomy

Data Pyramid for Embodied Manipulation

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-lan
HuggingFace Daily Papers 4d ago Research Agents & autonomy

A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-k content, but document relevance alone cannot localize, compose, or verify the evidence required by complex questions. Direct Corpus Interaction (DCI) enables such fine-grained operations through grep-style exploration, but its relevance-agnostic search can expose useful clues late and delay convergence. Recent advances use relevance to narrow
HuggingFace Daily Papers 4d ago Research Agents & autonomy
← Newer Older →