04:59 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility

Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augmented Generation, a deployed platform that makes the institutional knowledge at the Advanced Photon Source (APS) accessible to staff through natural-language queries, along with an operations-grounded
arXiv cs.AI 3d ago Research Agents & autonomy

Can Europe’s new AI safety regime tame US rogue agents — and Chinese ambitions?

The first-ever case of an artificial intelligence agent going rogue coincides with the EU's major new powers to regulate AI. But Europe's leverage may be limited by the U.S.-China two-way race for AI supremacy.
Politico Europe Technology 3d ago News RegulationSafety & alignment

Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents

Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent's context upon reading unvetted data, severely restricting downstream utility. We present APPA (Agentic Permissions Policy Algebra), an IFC framework that resolves this usability bottleneck through engine-managed context
arXiv cs.AI 3d ago Research RegulationPrivacy

Nvidia versammelt Tech-Schwergewichte für neue KI-Allianz

Mit einer breiten Branchenallianz will Nvidia KI-Agenten sicherer machen und positioniert sich zugleich gegen staatliche Beschränkungen offener Modelle.
Heise Online (DE) 3d ago News Agents & autonomy

Intapp bets on ‘firm AI’ with general release of Celeste

Intapp this month announced that its agentic ‘AI coworker’ Celeste, previewed in February, is now generally available. The listed company positions Celeste as part of a new category of ‘firm […] The post Intapp bets on ‘firm AI’ with general release of Celeste appeared first on Legal IT Insider .
Legal IT Insider 3d ago News Agents & autonomy

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair

Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branche
arXiv cs.AI 3d ago Research Agents & autonomy

☕️ OpenAI aurait mis une semaine à s’apercevoir que son agent avait attaqué Hugging Face

Le 21 juillet, OpenAI a publié un communiqué étonnant : un de ses systèmes IA était responsable de l’attaque orchestrée contre Hugging Face. Des détails étaient fournis, mais l’histoire gardait des zones d’ombre. Si l’incident a été transformé en opportunité commerciale, il semble le résultat d’une vaste carence en sécurité. Dans notre article du 22 juillet, […]
Next (FR, ex-INpact) 3d ago News Agents & autonomy

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent advances in Large Language Model (LLM) agents have enabled the automation of weather-related tasks, existing studies remain centered on isolated scientific tasks and overlook the chain of interdependent
arXiv cs.AI 3d ago Research Jobs & economyAgents & autonomy

Is open source the answer to rogue AI agents? Nvidia's new alliance says yes

As AI cybersecurity incidents ramp up, companies are racing to find a fix.
ZDNet AI 3d ago News Agents & autonomy

C’est parti : les poids de Kimi K3, plus gros modèle IA ouvert au monde, sont en ligne

Après des semaines de bruit autour de ses benchmarks et de ses ambitions agentiques, Moonshot AI a mis en ligne ce lundi les poids complets de Kimi K3 sur Hugging Face, ouvrant la voie à un déploiement et une adaptation par n'importe quel développeur.
Numerama (FR) 3d ago News Agents & autonomy

Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for automated model design, but when guided only by aggregate performance metrics, they lack insight into why individual cases fail and how the classifier should be revised. We present RecursiveECG, an evidence-driven LLM-as-Designer framework in which a
arXiv cs.AI 3d ago Research Agents & autonomy

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker

The warning shots will continue until civilization wakes up
Import AI 3d ago Field notes Agents & autonomy

FLUX 3: from video generation to robot control

The same multimodal backbone that generates 20-second video with audio is being tested on industrial manipulation at Audi Production Lab.
Air Street Capital (State of AI) 3d ago Field notes Agents & autonomy

Enigma raises $71M to make controlling a robot as easy as adjusting the volume

The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners.
TechCrunch 3d ago News Agents & autonomy

“Developers see this as the future”: Pilot Protocol launches to power the agent economy

When we created software agents, we built them in the shape of humans, as solitary individuals. Today, agents created by The post “Developers see this as the future”: Pilot Protocol launches to power the agent economy appeared first on The New Stack .
The New Stack AI 3d ago News Jobs & economyAgents & autonomy

Cloudflare open-sources a debugger for privacy protocols used by Apple and Microsoft, with AI agents in mind

Every time you use a privacy service like Apple’s iCloud Private Relay, you’re trusting a system designed so that no The post Cloudflare open-sources a debugger for privacy protocols used by Apple and Microsoft, with AI agents in mind appeared first on The New Stack .
The New Stack AI 3d ago News PrivacyAgents & autonomy

Dynatrace’s new agents can reveal the single hardest part of AI operations

Observability platform company Dynatrace announced a set of advancements to its Dynatrace Intelligence service on Monday that aim to take The post Dynatrace’s new agents can reveal the single hardest part of AI operations appeared first on The New Stack .
The New Stack AI 3d ago News Agents & autonomy

METR introduces a new metric to calculate exactly when AI agents become more expensive than humans

METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture. The article METR introduces a new metric to calculate exactly when AI agents become more expensive than humans appeared first on The Decoder .
The Decoder 3d ago News Agents & autonomy

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body but should not determine a recipient, account, command, or credential. Existing statistical methods typically control risk over the entire action, allowing failures in rare, high-risk fields to be obscured by benign arguments. We introduce role-stratified per-field conformal risk control, a calibration layer that wraps any per-field detector and sets
arXiv cs.AI 3d ago Research Agents & autonomy

Boss of startup hacked by rogue OpenAI agent urges ‘radical transparency’ in investigation

Artificial intelligence firm should provide $100m for cyber defences, says Hugging Face CEO The boss of the startup hacked by an OpenAI agent has called for the investigation into the incident to show “radical transparency”. Clément Delangue, the chief executive of Hugging Face, said the “unprecedented” attack on his business required a similar response. Continue reading...
The Guardian 3d ago News Military & securityAgents & autonomy

How would AI data centers in space even work? A former NASA robotics chief explains

'You collect electricity in space, and you eject heat in space. The only thing that comes to Earth is data.'
ZDNet AI 3d ago News Agents & autonomy

The path to artificial superintelligence

Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain. But they all have their own distinct knowledge and objectives. Today they can exchange data, but they are not yet able to actually coordinate…
MIT Technology Review 3d ago News HealthcareAgents & autonomy

Hackers used autonomous AI agent to spy on Thailand's finance ministry

Hackers used an autonomous artificial intelligence agent to carry out a cyber-espionage campaign against Thailand's Ministry of Finance, researchers discovered.
The Record (Recorded Future News) 3d ago News Military & securityAgents & autonomy

Building the enterprise environment for agentic AI

For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems. The platform best-suited to run agents is built with proper CPU capacity, resilient data access, policy-aware tool use, observability, memory management, and the…
MIT Technology Review 3d ago News RegulationAgents & autonomy

Article: An Evolutionary Architecture Pattern for Managing AI’s Pace of Change

Traditional API gateways assume deterministic services and simple schemas - assumptions agentic AI breaks. Discover why enterprise engineering leaders are adopting AI Gateways as an evolutionary architecture seam. Centralize guardrails, model routing, agent identity, action policy, and semantic audit within a single control plane to prevent costly incidents while keeping core platforms stable. By Joe Price, Branimir Đurek, Pavlos Migkiros, Trevor Dearham
InfoQ AI/ML 3d ago News RegulationAgents & autonomy

AgiBot starts Hong Kong IPO process

Chinese humanoid robot maker AgiBot has initiated the process for a Hong Kong IPO. The move follows earlier reports that the Shanghai-based company had selected banks and was preparing for a listing. Founded in 2023, AgiBot develops humanoid robots and embodied AI systems for industrial and commercial applications. The company has also released AgiBot World, […]
TechNode (CN) 3d ago News Agents & autonomyFinance, VC & PE

Surgical Re-enactment for Operating Room Workflow Datasets

The introduction of new technologies, such as surgical robots, is driving the vision of a connected, smart operating room (OR). However, realizing this vision requires a deep understanding of surgical workflows, which relies on realistic datasets capturing the actions of all OR personnel from both full room and surgical field perspectives. Acquiring such data in real ORs is prohibitively challenging due to factors such as ethics committee approvals, limited space for camera installation, and ste
arXiv cs.RO (robot ethics) 3d ago Research Agents & autonomy

NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics

Hugging Face Blog 3d ago Field notes Agents & autonomy

Vierbeiniger Roboter läuft über Wasser und soll Ertrinkende retten

Der QuadBoat-Roboter hat vier Beine, mit denen er übers Wasser laufen kann. Er soll bei Rettungsmissionen Ertrinkende bergen können.
Heise Online (DE) 3d ago News Agents & autonomy

Anzeige: KI-Agent erzeugt druckbares 3D-Modell in wenigen Minuten

Aus Text oder Foto entsteht schnell ein 3D-Mesh - doch das Generieren ist nur der Anfang. Das Modell muss noch optimiert, gepr�ft und druckbar gemacht werden. Der Meshy 3D-Agent b�ndelt diese verstreuten Schritte in einem einzigen Dialog. ( 3D-Drucker , KI )
Golem (DE) 3d ago News Agents & autonomy

AI agents: Augmenting, not replacing, human expertise in ICT professional services

Local customers are realising that AI and agentic AI are likely to be a complementary technology, augmenting the work humans do, says Paul Field, professional services/managed services BU leader at CASA Software.
ITWeb (ZA) 3d ago News Agents & autonomy

Les IA d’OpenAI qui dérapent, une faille Linux trouvée par Claude et un million de VM patchées d’urgence chez OVHcloud : on vous raconte la semaine Cyberguerre

Trois actualités à retenir cette semaine dans le cyberespace : des agents d'OpenAI qui ont piraté Hugging Face en toute autonomie, une faille noyau critique débusquée par Claude Mythos, et le récit d'une course contre la montre chez OVHcloud pour colmater une vulnérabilité vieille de 16 ans.
Numerama (FR) 3d ago News Agents & autonomy

Co-design of LLM-based preference agents: participation may drive overtrust

arXiv:2607.21757v1 Announce Type: new Abstract: Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designe
arXiv cs.CY 4d ago Research Agents & autonomy

From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE

arXiv:2607.21608v1 Announce Type: cross Abstract: With the EU AI Act entering into force, organizations developing or operating AI systems face new obligations on transparency, risk management, and traceability. For Requirements Engineering (RE), these obligations must be translated into testable, auditable requirements and verifiable evidence. However, many organizations currently lack systematic processes to achieve this. We hypothesize that LLM-based agentic validation tools can support this
arXiv cs.CY 4d ago Research RegulationAgents & autonomy

Agentic Evaluation of Copyright Law Compliance

arXiv:2607.21799v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate, reproducing that content. LLM agents should comply with the law, including copyright law. Presently, however, we lack adequate frameworks to assess whether they do so in practice. To that end, we introduce \textbf{Copyright-Bench}, a benchmark designed to evaluate \textit{LLM agents' compliance wi
arXiv cs.CY 4d ago Research RegulationCopyright & IP

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

arXiv:2505.19212v2 Announce Type: replace-cross Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and strategic behavior separately, there is limited understanding of how they act when moral imperatives directly conflict with profit incentives. We introduce \msimfull (\msim) to evaluate how LLMs behave in the priso
arXiv cs.CY 4d ago Research Safety & alignmentAgents & autonomy

Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

arXiv:2604.26577v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this context remains poorly characterized. We introduce a dataset of 270 harmful instructions spanning nine prohibited behavior categories grounded in the American Medical Association Principles of Medical Ethics, and use it to evaluate 72 LLMs in a simulation environment based on the Robotic H
arXiv cs.CY 4d ago Research HealthcareAgents & autonomy

BYD to unveil first humanoid robot in early August

BYD yesterday released a teaser poster hinting that its first humanoid robot will make its public debut in early August. The announcement follows earlier comments from BYD executives signaling the company’s ambitions in robotics. Last month, BYD Executive Vice President Li Ke said the company hopes to eventually deploy humanoid robots across its dealership network, […]
TechNode (CN) 4d ago News Agents & autonomy

Crowd navigation in a multi-room environment: a model predictive control framework for mobile robots

Mobile robots operating in human-populated environments must navigate complex, multi-room spaces while ensuring safety, i.e., generating collision-free motion. In this study, we present a sensor-based model predictive control (MPC) scheme designed for safe crowd navigation in such non-convex environments. The proposed framework decomposes the free space into a set of overlapping convex regions to construct a topological graph, enabling a high-level planner to compute optimal sequences of travers
Frontiers in Robotics and AI 4d ago Research Agents & autonomyEnvironment

Robotic lava tube mapping and multimodal data collection using quadruped and LiDAR

As part of TU Delft Rhizome 2.0 and Moonshot projects, focusing on the development of extraterrestrial habitats in lava tubes, the robotic mapping of an analogue lava tube in Sicily has been studied with the future goal of assessing its suitability for building construction. The main objective of the research was to survey a lava tube and acquire a novel dataset for future research, while also analyzing the collected data to evaluate possible future lava tube exploration scenarios for the Moon a
Frontiers in Robotics and AI 4d ago Research Agents & autonomy
← Newer Older →