Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility
Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augmented Generation, a deployed platform that makes the institutional knowledge at the Advanced Photon Source (APS) accessible to staff through natural-language queries, along with an operations-grounded
Can Europe’s new AI safety regime tame US rogue agents — and Chinese ambitions?
The first-ever case of an artificial intelligence agent going rogue coincides with the EU's major new powers to regulate AI. But Europe's leverage may be limited by the U.S.-China two-way race for AI supremacy.
Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent's context upon reading unvetted data, severely restricting downstream utility. We present APPA (Agentic Permissions Policy Algebra), an IFC framework that resolves this usability bottleneck through engine-managed context
Nvidia versammelt Tech-Schwergewichte für neue KI-Allianz
Mit einer breiten Branchenallianz will Nvidia KI-Agenten sicherer machen und positioniert sich zugleich gegen staatliche Beschränkungen offener Modelle.
Intapp bets on ‘firm AI’ with general release of Celeste
Intapp this month announced that its agentic ‘AI coworker’ Celeste, previewed in February, is now generally available. The listed company positions Celeste as part of a new category of ‘firm […] The post Intapp bets on ‘firm AI’ with general release of Celeste appeared first on Legal IT Insider .
Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branche
☕️ OpenAI aurait mis une semaine à s’apercevoir que son agent avait attaqué Hugging Face
Le 21 juillet, OpenAI a publié un communiqué étonnant : un de ses systèmes IA était responsable de l’attaque orchestrée contre Hugging Face. Des détails étaient fournis, mais l’histoire gardait des zones d’ombre. Si l’incident a été transformé en opportunité commerciale, il semble le résultat d’une vaste carence en sécurité. Dans notre article du 22 juillet, […]
SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents
Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent advances in Large Language Model (LLM) agents have enabled the automation of weather-related tasks, existing studies remain centered on isolated scientific tasks and overlook the chain of interdependent
Is open source the answer to rogue AI agents? Nvidia's new alliance says yes
As AI cybersecurity incidents ramp up, companies are racing to find a fix.
C’est parti : les poids de Kimi K3, plus gros modèle IA ouvert au monde, sont en ligne
Après des semaines de bruit autour de ses benchmarks et de ses ambitions agentiques, Moonshot AI a mis en ligne ce lundi les poids complets de Kimi K3 sur Hugging Face, ouvrant la voie à un déploiement et une adaptation par n'importe quel développeur.
Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers
Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for automated model design, but when guided only by aggregate performance metrics, they lack insight into why individual cases fail and how the classifier should be revised. We present RecursiveECG, an evidence-driven LLM-as-Designer framework in which a
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker
The warning shots will continue until civilization wakes up
FLUX 3: from video generation to robot control
The same multimodal backbone that generates 20-second video with audio is being tested on industrial manipulation at Audi Production Lab.
Enigma raises $71M to make controlling a robot as easy as adjusting the volume
The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners.
“Developers see this as the future”: Pilot Protocol launches to power the agent economy
When we created software agents, we built them in the shape of humans, as solitary individuals. Today, agents created by The post “Developers see this as the future”: Pilot Protocol launches to power the agent economy appeared first on The New Stack .
Cloudflare open-sources a debugger for privacy protocols used by Apple and Microsoft, with AI agents in mind
Every time you use a privacy service like Apple’s iCloud Private Relay, you’re trusting a system designed so that no The post Cloudflare open-sources a debugger for privacy protocols used by Apple and Microsoft, with AI agents in mind appeared first on The New Stack .
Dynatrace’s new agents can reveal the single hardest part of AI operations
Observability platform company Dynatrace announced a set of advancements to its Dynatrace Intelligence service on Monday that aim to take The post Dynatrace’s new agents can reveal the single hardest part of AI operations appeared first on The New Stack .
METR introduces a new metric to calculate exactly when AI agents become more expensive than humans
METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture. The article METR introduces a new metric to calculate exactly when AI agents become more expensive than humans appeared first on The Decoder .
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body but should not determine a recipient, account, command, or credential. Existing statistical methods typically control risk over the entire action, allowing failures in rare, high-risk fields to be obscured by benign arguments. We introduce role-stratified per-field conformal risk control, a calibration layer that wraps any per-field detector and sets
Boss of startup hacked by rogue OpenAI agent urges ‘radical transparency’ in investigation
Artificial intelligence firm should provide $100m for cyber defences, says Hugging Face CEO The boss of the startup hacked by an OpenAI agent has called for the investigation into the incident to show “radical transparency”. Clément Delangue, the chief executive of Hugging Face, said the “unprecedented” attack on his business required a similar response. Continue reading...
How would AI data centers in space even work? A former NASA robotics chief explains
'You collect electricity in space, and you eject heat in space. The only thing that comes to Earth is data.'
The path to artificial superintelligence
Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain. But they all have their own distinct knowledge and objectives. Today they can exchange data, but they are not yet able to actually coordinate…
Hackers used autonomous AI agent to spy on Thailand's finance ministry
Hackers used an autonomous artificial intelligence agent to carry out a cyber-espionage campaign against Thailand's Ministry of Finance, researchers discovered.
Building the enterprise environment for agentic AI
For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems. The platform best-suited to run agents is built with proper CPU capacity, resilient data access, policy-aware tool use, observability, memory management, and the…
Article: An Evolutionary Architecture Pattern for Managing AI’s Pace of Change
Traditional API gateways assume deterministic services and simple schemas - assumptions agentic AI breaks. Discover why enterprise engineering leaders are adopting AI Gateways as an evolutionary architecture seam. Centralize guardrails, model routing, agent identity, action policy, and semantic audit within a single control plane to prevent costly incidents while keeping core platforms stable. By Joe Price, Branimir Đurek, Pavlos Migkiros, Trevor Dearham
AgiBot starts Hong Kong IPO process
Chinese humanoid robot maker AgiBot has initiated the process for a Hong Kong IPO. The move follows earlier reports that the Shanghai-based company had selected banks and was preparing for a listing. Founded in 2023, AgiBot develops humanoid robots and embodied AI systems for industrial and commercial applications. The company has also released AgiBot World, […]
Surgical Re-enactment for Operating Room Workflow Datasets
The introduction of new technologies, such as surgical robots, is driving the vision of a connected, smart operating room (OR). However, realizing this vision requires a deep understanding of surgical workflows, which relies on realistic datasets capturing the actions of all OR personnel from both full room and surgical field perspectives. Acquiring such data in real ORs is prohibitively challenging due to factors such as ethics committee approvals, limited space for camera installation, and ste
NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics
Vierbeiniger Roboter läuft über Wasser und soll Ertrinkende retten
Der QuadBoat-Roboter hat vier Beine, mit denen er übers Wasser laufen kann. Er soll bei Rettungsmissionen Ertrinkende bergen können.
Anzeige: KI-Agent erzeugt druckbares 3D-Modell in wenigen Minuten
Aus Text oder Foto entsteht schnell ein 3D-Mesh - doch das Generieren ist nur der Anfang. Das Modell muss noch optimiert, gepr�ft und druckbar gemacht werden. Der Meshy 3D-Agent b�ndelt diese verstreuten Schritte in einem einzigen Dialog. ( 3D-Drucker , KI )
AI agents: Augmenting, not replacing, human expertise in ICT professional services
Local customers are realising that AI and agentic AI are likely to be a complementary technology, augmenting the work humans do, says Paul Field, professional services/managed services BU leader at CASA Software.
Les IA d’OpenAI qui dérapent, une faille Linux trouvée par Claude et un million de VM patchées d’urgence chez OVHcloud : on vous raconte la semaine Cyberguerre
Trois actualités à retenir cette semaine dans le cyberespace : des agents d'OpenAI qui ont piraté Hugging Face en toute autonomie, une faille noyau critique débusquée par Claude Mythos, et le récit d'une course contre la montre chez OVHcloud pour colmater une vulnérabilité vieille de 16 ans.
Co-design of LLM-based preference agents: participation may drive overtrust
arXiv:2607.21757v1 Announce Type: new Abstract: Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designe
From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE
arXiv:2607.21608v1 Announce Type: cross Abstract: With the EU AI Act entering into force, organizations developing or operating AI systems face new obligations on transparency, risk management, and traceability. For Requirements Engineering (RE), these obligations must be translated into testable, auditable requirements and verifiable evidence. However, many organizations currently lack systematic processes to achieve this. We hypothesize that LLM-based agentic validation tools can support this
Agentic Evaluation of Copyright Law Compliance
arXiv:2607.21799v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate, reproducing that content. LLM agents should comply with the law, including copyright law. Presently, however, we lack adequate frameworks to assess whether they do so in practice. To that end, we introduce \textbf{Copyright-Bench}, a benchmark designed to evaluate \textit{LLM agents' compliance wi
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
arXiv:2505.19212v2 Announce Type: replace-cross Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and strategic behavior separately, there is limited understanding of how they act when moral imperatives directly conflict with profit incentives. We introduce \msimfull (\msim) to evaluate how LLMs behave in the priso
Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
arXiv:2604.26577v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this context remains poorly characterized. We introduce a dataset of 270 harmful instructions spanning nine prohibited behavior categories grounded in the American Medical Association Principles of Medical Ethics, and use it to evaluate 72 LLMs in a simulation environment based on the Robotic H
BYD to unveil first humanoid robot in early August
BYD yesterday released a teaser poster hinting that its first humanoid robot will make its public debut in early August. The announcement follows earlier comments from BYD executives signaling the company’s ambitions in robotics. Last month, BYD Executive Vice President Li Ke said the company hopes to eventually deploy humanoid robots across its dealership network, […]
Crowd navigation in a multi-room environment: a model predictive control framework for mobile robots
Mobile robots operating in human-populated environments must navigate complex, multi-room spaces while ensuring safety, i.e., generating collision-free motion. In this study, we present a sensor-based model predictive control (MPC) scheme designed for safe crowd navigation in such non-convex environments. The proposed framework decomposes the free space into a set of overlapping convex regions to construct a topological graph, enabling a high-level planner to compute optimal sequences of travers
Robotic lava tube mapping and multimodal data collection using quadruped and LiDAR
As part of TU Delft Rhizome 2.0 and Moonshot projects, focusing on the development of extraterrestrial habitats in lava tubes, the robotic mapping of an analogue lava tube in Sicily has been studied with the future goal of assessing its suitability for building construction. The main objective of the research was to survey a lava tube and acquire a novel dataset for future research, while also analyzing the collected data to evaluate possible future lava tube exploration scenarios for the Moon a