Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
What's Exciting in Payments Today?
Joining the FinextraTV studio at Payments Unleashed in London, David Budzevski, Global Acceptance, Mastercard helped to discuss how payments is evolving and what it means for organisations operating in the industry. Budzevski outlined the changes in AI, and emphasised the issue of trust - something that he says will continue to define what makes a company competitive as AI scales. He explained how agentic commerce entirely changes the typical relationships that consumers and leaders alike have g
Datacenter Capex is Spilling over into a ChatGPT of Robotics Moment set for 2027 and this decade.
Do people really want datacenters, robots and AI overlords? This is going to become a problem. The robotics flood is near. 🤖
Article: Multi-Agent AI for Production Security Operations: An A2A and MCP Architecture in a 5G Core
The bottleneck in a mature SOC is rarely analyst triage; rather, it is the detection-engineering team's ability to keep the rule base aligned with a threat landscape that evolves faster than rules can be written. Learn how multi-agent system for production security operations has reduced mean times to detect and to respond by 40% and compressed the human work required by 12x. By Willem Berroubache
Microsoft Agent Framework bietet ein Harness für .NET und Python
Das Agent Framework bringt nun ein vorgefertigtes Harness mit, um aus großen Sprachmodellen Agenten zu machen.
Flank Launches ‘Record’ – Agentic Contract Truth System
Flank, the agentic legal tech company focused on inhouse teams, has launched Flank Record, an autonomous agentic contract system of record, which greatly reduces the ...
HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices
Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited personalisation. The emergence of LLM agents creates a new opportunity for Personal Health Agentic Analysis, where health insights can be generated adaptively and in context. However, currently there is no open-source locally deployable platform capable of processing personal health data in real time while preserving privacy. We present HiMe, a locally deplo
Alibaba unit says 5-in-1 AI gives robots unified brain, body and limbs
Alibaba Group Holding’s mapping unit has unveiled an upgrade to what it calls the world’s first technology framework that unites a robot’s “feet, hands, brains, central nerves and motor nerves” into a single system. Amap’s ABot system marks the latest move in Chinese tech players’ race to harness artificial intelligence models to make robots more capable. Referred to as embodied intelligence, the effort aims to equip machines with the “brains” and other elements required to navigate and complete
The robot byline is quietly disappearing
The most obvious use of generative AI is writing. It’s right there in the name—large language models (LLMs) are all about reading, organizing, analyzing, and conjuring words—which is exactly why many in the journalism profession have been going through a kind of existential crisis these past few years. And the crisis isn’t just theoretical. As artificial intelligence systems get better at writing, a growing number of newsrooms are using AI to help not just with analysis, process, and ideas, but
heise-Angebot: iX-Workshop: Claude Code in der Praxis – effizienter entwickeln mit KI-Agenten
Erfahren Sie, wie Sie Ihre Entwicklungsaufgaben mit Claude Code autonom bearbeiten lassen und Ihre Workflows mit KI-Agenten spürbar beschleunigen können.
Assured Health Secures $19M to Get Providers In-Network Faster with Agentic AI
Assured Health raised $19 million for its AI agents that verify provider credentials and manage insurance enrollment for health systems and group practices. The startup said it can cut a process that usually takes months down to days. The post Assured Health Secures $19M to Get Providers In-Network Faster with Agentic AI appeared first on MedCity News .
Clinical Pathways as Safety Specifications for Physical AI in Hospital Wards
arXiv:2607.19827v1 Announce Type: cross Abstract: Ensuring safety in Physical AI systems operating in real-world environments is a critical challenge, particularly in hospital wards where vulnerable patients, clinical staff, medical devices, and assistive robots coexist. In this paper, we reinterpret Clinical Pathways as explicit runtime safety specifications for embodied medical AI. We propose a conceptual robotic architecture that integrates wearable sensors, smart medical devices, and assisti
When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets
arXiv:2607.19967v1 Announce Type: cross Abstract: Shippers are beginning to delegate carrier selection to large language model (LLM) agents. We ask what such delegation does to a freight matching market, and which platform design choices contain it. We carried out agent-based simulations in which fifty shipper agents, built on commercial LLMs from OpenAI (GPT), Anthropic (Claude), and Google (Gemini), procure truckload capacity for thirty days. The market implements the rules of digital freight
From Assistance to Autonomy -- A Researcher Study on the Potential of AI Support for Qualitative Data Analysis
arXiv:2501.19275v4 Announce Type: replace Abstract: The advent of AI technologies, such as Large Language Models, has introduced new possibilities for Qualitative Data Analysis (QDA), offering both opportunities and challenges. To help navigate the responsible integration of AI into QDA, we conducted semi-structured interviews with 15 Human-Computer Interaction (HCI) researchers experienced in QDA. While our participants were open to AI support in their QDA workflows, they expressed concerns abo
‘The Wild Robot’ VFX Studio STIM Sets Canary Islands Operation, Planning to Create More Than 100 Jobs
French animation and visual effects company STIM – whose credits include “The Wild Robot,” “The Garfield Movie” and “Coyote vs. Acme” – is opening a production center in Spain’s Canary Islands, with plans to create more than 100 jobs as it expands its European operations. STIM Tenerife is set to begin operating in August and […]
Tesla’s profits plunge as Musk’s carmaker uses discounts to boost sales
Capital expenditures more than doubled as company accelerates pivot to AI and robots
ServiceNow CEO defends the company's relevancy, touting a kill switch for rogue AI agents
ServiceNow CEO Bill McDermott said the company is becoming increasingly important as businesses deploy more autonomous AI agents.
The Fourth Circuit Says Border Agents Can Search Your Phone By Hand, No Suspicion Required
Legal intern Suzanne Castillo was the principal author of this post. The Fourth Circuit issued a disappointing opinion in U.S. v. Belmonte Cardozo , a case in which EFF filed an amicus brief , alongside the national ACLU, its Maryland, North Carolina, South Carolina, and Virginia affiliates, and the National Association of Criminal Defense Lawyers (NACDL). We argued that electronic device searches at the border should require a warrant based on probable cause, but at minimum, regardless of wheth
Energy Department demos Genesis Mission platform
Around 278 teams will gain access to the platform, including agentic frameworks and high-performance computing, among other resources. The post Energy Department demos Genesis Mission platform appeared first on FedScoop .
OpenAI’s rogue agents are a wake-up call to risks posed by artificial intelligence | Shakeel Hashim
Hacking of Hugging Face shows we do not seem to have reliable ways to curb extremely powerful AI systems Last week Hugging Face – a company that hosts artificial intelligence models and datasets – was hacked . After it reported the incident to law enforcement, few would have predicted what came next: the culprits were revealed to be AI agents from OpenAI, which had broken out of containment and were acting of their own accord. Shakeel Hashim is the editor of Transformer , a publication about the
OpenForgeRL: Train Harness-native Agents in Any Environment
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments.
Sample-Efficient Learning from Agent Experience
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is removed from the context. Separately, context distillation provides a mechanism for internalizing contextual information into model weights. However, applying it to agents' interacti
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-gener
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce ARE
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and re
TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation
The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout hallucination or simplified procedural generation, they frequently suffer from physical implausibility and fail to capture the complex, dense clutter of actual human environments. In this paper, we introduce TableVerse, a fully automated Real2Sim pipeline that shift
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registe
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities including planning, requirement clarification, tool use, debugging, and repository-level construction. Yet existing benchmarks have not fully caught up with this shift, evaluating agents on static, fully spec
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing i
Northrop Grumman launches in-space servicing satellites for life-extension missions
The Mission Robotic Vehicle will attach life-extending orbital "jetpacks" to sustain three commercial satellites years after they run out of fuel. The post Northrop Grumman launches in-space servicing satellites for life-extension missions appeared first on DefenseScoop .
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.
Travis Kalanick’s robotics company raises $1.7B, led by a16z
Uber is also investing in Travis Kalanick's company Atoms, which has made gauzy claims about using industrial AI to modernize the world.
OpenAI cyber models broke out of training environment to hack Hugging Face
The incident is unique because it was "driven, end to end, by an autonomous AI agent system," according to Hugging Face.
The Real Lesson of OpenAI's 'Rogue' Agent Isn't Alignment
Hyundai claims humanoid robot plan is not part of talks with striking workers
Union previously warned automaker that any robot deployment must be negotiated.
This is the stock to buy after OpenAI's AI agent goes rogue in a cybersecurity test
Jim Cramer said CrowdStrike is the stock to buy after OpenAI's agentic breach.
Agents keep changing their answers. Harness just built delivery pipelines that don’t care.
Software delivery lifecycle company (SDLC) Harness wants to put agents through the same pipelines and controls that all application code The post Agents keep changing their answers. Harness just built delivery pipelines that don’t care. appeared first on The New Stack .
Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement Learning
Full-sized humanoid robot capabilities have grown exponentially in recent years, aiming towards general-purpose deployment in human environments. A popular control method used by manufacturers utilizes Virtual Reality for upper-body teleoperation and Reinforcement Learning for lower-body balance and locomotion control. As a result, a single remote operator can see, manipulate, and navigate about a real, distant physical environment. This powerful control stack is often relegated to expensive ful
OpenAI built support agents for its own customer service line, now it hopes big enterprises will trust them too
The general consensus emerging across the AI and industrial spheres is that the models themselves are no longer the bottleneck The post OpenAI built support agents for its own customer service line, now it hopes big enterprises will trust them too appeared first on The New Stack .
I gave Perplexity's agentic AI 5 complex tasks to run on my Mac - and I'll do it again
Perplexity's Mac app offers its own agentic AI, Personal Computer, which can handle multi-step tasks on your computer from start to finish. See why the results impressed me.
China lleva años perfeccionando robots que se pegan a las paredes con un objetivo: que nadie más muera colgado de un andamio
Durante estas semanas ha circulado un vídeo sobre un robot trepando por la superficie curva de un tanque metálico, pegado a la chapa como si fuera una lapa, mientras pule y elimina óxido sin que ningún humano tenga que colgarse de una cuerda para hacerlo. En nuestros lares es bastante sorprendente ver esto, pero en China hay un buen puñado de empresas que ofrecen robots especializados en trabajos de altura, y por eso se me ha ocurrido que quizás es buena idea contar en este artículo qué hay de i