12:31 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

What's Exciting in Payments Today?

Joining the FinextraTV studio at Payments Unleashed in London, David Budzevski, Global Acceptance, Mastercard helped to discuss how payments is evolving and what it means for organisations operating in the industry. Budzevski outlined the changes in AI, and emphasised the issue of trust - something that he says will continue to define what makes a company competitive as AI scales. He explained how agentic commerce entirely changes the typical relationships that consumers and leaders alike have g
Finextra AI 8d ago News Agents & autonomy

Datacenter Capex is Spilling over into a ChatGPT of Robotics Moment set for 2027 and this decade.

Do people really want datacenters, robots and AI overlords? This is going to become a problem. The robotics flood is near. 🤖
AI Supremacy 8d ago Field notes Agents & autonomyEnvironment

Article: Multi-Agent AI for Production Security Operations: An A2A and MCP Architecture in a 5G Core

The bottleneck in a mature SOC is rarely analyst triage; rather, it is the detection-engineering team's ability to keep the rule base aligned with a threat landscape that evolves faster than rules can be written. Learn how multi-agent system for production security operations has reduced mean times to detect and to respond by 40% and compressed the human work required by 12x. By Willem Berroubache
InfoQ AI/ML 8d ago News Agents & autonomy

Microsoft Agent Framework bietet ein Harness für .NET und Python

Das Agent Framework bringt nun ein vorgefertigtes Harness mit, um aus großen Sprachmodellen Agenten zu machen.
Heise Online (DE) 8d ago News Agents & autonomy

Flank Launches ‘Record’ – Agentic Contract Truth System

Flank, the agentic legal tech company focused on inhouse teams, has launched Flank Record, an autonomous agentic contract system of record, which greatly reduces the ...
Artificial Lawyer 8d ago News Agents & autonomy

HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices

Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited personalisation. The emergence of LLM agents creates a new opportunity for Personal Health Agentic Analysis, where health insights can be generated adaptively and in context. However, currently there is no open-source locally deployable platform capable of processing personal health data in real time while preserving privacy. We present HiMe, a locally deplo
arXiv cs.HC 8d ago Research PrivacyHealthcare

Alibaba unit says 5-in-1 AI gives robots unified brain, body and limbs

Alibaba Group Holding’s mapping unit has unveiled an upgrade to what it calls the world’s first technology framework that unites a robot’s “feet, hands, brains, central nerves and motor nerves” into a single system. Amap’s ABot system marks the latest move in Chinese tech players’ race to harness artificial intelligence models to make robots more capable. Referred to as embodied intelligence, the effort aims to equip machines with the “brains” and other elements required to navigate and complete
SCMP Tech (HK/CN) 8d ago News Agents & autonomy

The robot byline is quietly disappearing

The most obvious use of generative AI is writing. It’s right there in the name—large language models (LLMs) are all about reading, organizing, analyzing, and conjuring words—which is exactly why many in the journalism profession have been going through a kind of existential crisis these past few years. And the crisis isn’t just theoretical. As artificial intelligence systems get better at writing, a growing number of newsrooms are using AI to help not just with analysis, process, and ideas, but
Fast Company Tech 8d ago News Safety & alignmentAgents & autonomy

heise-Angebot: iX-Workshop: Claude Code in der Praxis – effizienter entwickeln mit KI-Agenten

Erfahren Sie, wie Sie Ihre Entwicklungsaufgaben mit Claude Code autonom bearbeiten lassen und Ihre Workflows mit KI-Agenten spürbar beschleunigen können.
Heise Online (DE) 8d ago News Agents & autonomy

Assured Health Secures $19M to Get Providers In-Network Faster with Agentic AI

Assured Health raised $19 million for its AI agents that verify provider credentials and manage insurance enrollment for health systems and group practices. The startup said it can cut a process that usually takes months down to days. The post Assured Health Secures $19M to Get Providers In-Network Faster with Agentic AI appeared first on MedCity News .
MedCity News AI 8d ago News HealthcareAgents & autonomy

Clinical Pathways as Safety Specifications for Physical AI in Hospital Wards

arXiv:2607.19827v1 Announce Type: cross Abstract: Ensuring safety in Physical AI systems operating in real-world environments is a critical challenge, particularly in hospital wards where vulnerable patients, clinical staff, medical devices, and assistive robots coexist. In this paper, we reinterpret Clinical Pathways as explicit runtime safety specifications for embodied medical AI. We propose a conceptual robotic architecture that integrates wearable sensors, smart medical devices, and assisti
arXiv cs.CY 8d ago Research HealthcareAgents & autonomy

When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets

arXiv:2607.19967v1 Announce Type: cross Abstract: Shippers are beginning to delegate carrier selection to large language model (LLM) agents. We ask what such delegation does to a freight matching market, and which platform design choices contain it. We carried out agent-based simulations in which fifty shipper agents, built on commercial LLMs from OpenAI (GPT), Anthropic (Claude), and Google (Gemini), procure truckload capacity for thirty days. The market implements the rules of digital freight
arXiv cs.CY 8d ago Research Agents & autonomy

From Assistance to Autonomy -- A Researcher Study on the Potential of AI Support for Qualitative Data Analysis

arXiv:2501.19275v4 Announce Type: replace Abstract: The advent of AI technologies, such as Large Language Models, has introduced new possibilities for Qualitative Data Analysis (QDA), offering both opportunities and challenges. To help navigate the responsible integration of AI into QDA, we conducted semi-structured interviews with 15 Human-Computer Interaction (HCI) researchers experienced in QDA. While our participants were open to AI support in their QDA workflows, they expressed concerns abo
arXiv cs.CY 8d ago Research Agents & autonomy

‘The Wild Robot’ VFX Studio STIM Sets Canary Islands Operation, Planning to Create More Than 100 Jobs

French animation and visual effects company STIM – whose credits include “The Wild Robot,” “The Garfield Movie” and “Coyote vs. Acme” – is opening a production center in Spain’s Canary Islands, with plans to create more than 100 jobs as it expands its European operations. STIM Tenerife is set to begin operating in August and […]
Variety (AI) 8d ago News Jobs & economyAgents & autonomy

Tesla’s profits plunge as Musk’s carmaker uses discounts to boost sales

Capital expenditures more than doubled as company accelerates pivot to AI and robots
Financial Times Technology (headlines) 8d ago News Agents & autonomy

ServiceNow CEO defends the company's relevancy, touting a kill switch for rogue AI agents

ServiceNow CEO Bill McDermott said the company is becoming increasingly important as businesses deploy more autonomous AI agents.
CNBC Technology 8d ago News RegulationAgents & autonomy

The Fourth Circuit Says Border Agents Can Search Your Phone By Hand, No Suspicion Required

Legal intern Suzanne Castillo was the principal author of this post. The Fourth Circuit issued a disappointing opinion in U.S. v. Belmonte Cardozo , a case in which EFF filed an amicus brief , alongside the national ACLU, its Maryland, North Carolina, South Carolina, and Virginia affiliates, and the National Association of Criminal Defense Lawyers (NACDL). We argued that electronic device searches at the border should require a warrant based on probable cause, but at minimum, regardless of wheth
EFF Deeplinks 8d ago Field notes Military & securityAgents & autonomy

Energy Department demos Genesis Mission platform

Around 278 teams will gain access to the platform, including agentic frameworks and high-performance computing, among other resources. The post Energy Department demos Genesis Mission platform appeared first on FedScoop .
FedScoop 8d ago News Agents & autonomyEnvironment

OpenAI’s rogue agents are a wake-up call to risks posed by artificial intelligence | Shakeel Hashim

Hacking of Hugging Face shows we do not seem to have reliable ways to curb extremely powerful AI systems Last week Hugging Face – a company that hosts artificial intelligence models and datasets – was hacked . After it reported the incident to law enforcement, few would have predicted what came next: the culprits were revealed to be AI agents from OpenAI, which had broken out of containment and were acting of their own accord. Shakeel Hashim is the editor of Transformer , a publication about the
The Guardian 8d ago News RegulationAgents & autonomy

OpenForgeRL: Train Harness-native Agents in Any Environment

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments.
HuggingFace Daily Papers 8d ago Research Agents & autonomyEnvironment

Sample-Efficient Learning from Agent Experience

Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is removed from the context. Separately, context distillation provides a mechanism for internalizing contextual information into model weights. However, applying it to agents' interacti
HuggingFace Daily Papers 8d ago Research Agents & autonomyEnvironment

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-gener
HuggingFace Daily Papers 8d ago Research Agents & autonomy

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce ARE
HuggingFace Daily Papers 8d ago Research Agents & autonomy

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and re
HuggingFace Daily Papers 8d ago Research Agents & autonomy

TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation

The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout hallucination or simplified procedural generation, they frequently suffer from physical implausibility and fail to capture the complex, dense clutter of actual human environments. In this paper, we introduce TableVerse, a fully automated Real2Sim pipeline that shift
HuggingFace Daily Papers 8d ago Research Agents & autonomyEnvironment

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registe
HuggingFace Daily Papers 8d ago Research Agents & autonomy

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities including planning, requirement clarification, tool use, debugging, and repository-level construction. Yet existing benchmarks have not fully caught up with this shift, evaluating agents on static, fully spec
HuggingFace Daily Papers 8d ago Research Agents & autonomy

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing i
HuggingFace Daily Papers 8d ago Research Agents & autonomy

Northrop Grumman launches in-space servicing satellites for life-extension missions

The Mission Robotic Vehicle will attach life-extending orbital "jetpacks" to sustain three commercial satellites years after they run out of fuel. The post Northrop Grumman launches in-space servicing satellites for life-extension missions appeared first on DefenseScoop .
DefenseScoop 8d ago News Agents & autonomy

OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation

This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.
Dont Worry About the Vase (Zvi) 8d ago Field notes Agents & autonomy

Travis Kalanick’s robotics company raises $1.7B, led by a16z

Uber is also investing in Travis Kalanick's company Atoms, which has made gauzy claims about using industrial AI to modernize the world.
TechCrunch 8d ago News Agents & autonomyFinance, VC & PE

OpenAI cyber models broke out of training environment to hack Hugging Face

The incident is unique because it was "driven, end to end, by an autonomous AI agent system," according to Hugging Face.
CNBC Technology 8d ago News Military & securityAgents & autonomy

The Real Lesson of OpenAI's 'Rogue' Agent Isn't Alignment

Tech Policy Press 8d ago News Safety & alignmentAgents & autonomy

Hyundai claims humanoid robot plan is not part of talks with striking workers

Union previously warned automaker that any robot deployment must be negotiated.
Ars Technica 8d ago News Agents & autonomy

This is the stock to buy after OpenAI's AI agent goes rogue in a cybersecurity test

Jim Cramer said CrowdStrike is the stock to buy after OpenAI's agentic breach.
CNBC Technology 8d ago News Agents & autonomy

Agents keep changing their answers. Harness just built delivery pipelines that don’t care.

Software delivery lifecycle company (SDLC) Harness wants to put agents through the same pipelines and controls that all application code The post Agents keep changing their answers. Harness just built delivery pipelines that don’t care. appeared first on The New Stack .
The New Stack AI 8d ago News Agents & autonomy

Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement Learning

Full-sized humanoid robot capabilities have grown exponentially in recent years, aiming towards general-purpose deployment in human environments. A popular control method used by manufacturers utilizes Virtual Reality for upper-body teleoperation and Reinforcement Learning for lower-body balance and locomotion control. As a result, a single remote operator can see, manipulate, and navigate about a real, distant physical environment. This powerful control stack is often relegated to expensive ful
arXiv cs.HC 8d ago Research Agents & autonomyEnvironment

OpenAI built support agents for its own customer service line, now it hopes big enterprises will trust them too

The general consensus emerging across the AI and industrial spheres is that the models themselves are no longer the bottleneck The post OpenAI built support agents for its own customer service line, now it hopes big enterprises will trust them too appeared first on The New Stack .
The New Stack AI 8d ago News Agents & autonomy

I gave Perplexity's agentic AI 5 complex tasks to run on my Mac - and I'll do it again

Perplexity's Mac app offers its own agentic AI, Personal Computer, which can handle multi-step tasks on your computer from start to finish. See why the results impressed me.
ZDNet AI 8d ago News Agents & autonomy

China lleva años perfeccionando robots que se pegan a las paredes con un objetivo: que nadie más muera colgado de un andamio

Durante estas semanas ha circulado un vídeo sobre un robot trepando por la superficie curva de un tanque metálico, pegado a la chapa como si fuera una lapa, mientras pule y elimina óxido sin que ningún humano tenga que colgarse de una cuerda para hacerlo. En nuestros lares es bastante sorprendente ver esto, pero en China hay un buen puñado de empresas que ofrecen robots especializados en trabajos de altura, y por eso se me ha ocurrido que quizás es buena idea contar en este artículo qué hay de i
Xataka (ES) 8d ago News Agents & autonomy
← Newer Older →