23:20 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

Medical robotics beyond automation: Human-robot collaboration and the RONNA system as a socio-technical case study

Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Marina Raguž, Domagoj Dlaka, Marko Švaco, Petar Marčinković, Dominik Romić, Filip Šuligoj, Bojan Šekoranja, Darko Chudy, Bojan Jerbić
Technology in Society 18h ago Research Jobs & economyHealthcare

A scoping review of generative AI-powered agentic AI in education: Research landscape, agentic capabilities, and insights from the frontier agent paradigm, exemplified by OpenClaw

Publication date: Available online 28 July 2026 Source: Computers and Education: Artificial Intelligence Author(s): Ningxia Wang, Di Zou, Haoran Xie, S.Joe Qin
Computers and Education: Artificial Intelligence 18h ago Research Children & educationAgents & autonomy

OpenAIs KI-Agent knackte nicht nur Hugging Face – was genau passierte

OpenAI-Modelle attackierten vergangene Woche scheinbar weitere Software. Die Firma erklärt nun ausführlicher, was genau passiert ist.
Heise Online (DE) 18h ago News Agents & autonomy

The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

arXiv:2607.26064v1 Announce Type: new Abstract: AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversight. As seen by increased submissions to ML venues, the verification gap between scientific output and our ability to check it is already widening, and autonomous agents make it worse by magnitudes given human-agent asymmetry. We argue that science must evolve its verification infrastructure, as it ha
arXiv cs.CY 19h ago Research Agents & autonomy

"Nobody Did This": Contribution, Originality, and Accountability in Agent-Mediated Collaboration

arXiv:2607.26387v1 Announce Type: new Abstract: Collaborative knowledge work is changing in ways that go beyond disclosure or transparency. LLM agents are now embedded in how teams research, design, write, and decide: mediating between members, synthesizing inputs, reformulating ideas, and drafting shared outputs. They do not only facilitate collaboration; they operate within the workflow at the moment contributions are being formed. In doing so, they risk undermining the social conditions under
arXiv cs.CY 19h ago Research Agents & autonomyTransparency

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

arXiv:2607.26313v1 Announce Type: cross Abstract: Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are metadata-borne: a stale price or a superseded record, perfectly well-formed in the payload and betrayed only by freshness, lineage, or provenance. Such a defect never enters the agent's context, and an agent cannot doubt data it cannot see. On a priced replenishment benchmark, a competent agent sile
arXiv cs.CY 19h ago Research Agents & autonomy

Can AI agents conduct open-ended AI research? Early evidence from two case studies

arXiv:2607.27191v1 Announce Type: cross Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D auto
arXiv cs.CY 19h ago Research Agents & autonomy

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

arXiv:2604.24155v4 Announce Type: replace Abstract: The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically. This paper extends that challenge by examining tw
arXiv cs.CY 19h ago Research Safety & alignmentAgents & autonomy

The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning

arXiv:2507.04398v3 Announce Type: replace-cross Abstract: Generative AI is becoming part of academic writing, but its educational value depends on how control is shared between learner and system. This study examined an agency gap: performance differences that may arise when AI agent initiative is misaligned with learners' generative AI literacy. Seventy-nine medical and nursing students completed two multimodal analytical writing tasks using healthcare simulation data visualisations. They were
arXiv cs.CY 19h ago Research Safety & alignmentHealthcare

China threatens retaliation against U.S. humanoid robot ban, says it 'severely damages' relations

China's Commerce Ministry said Thursday the U.S. Federal Communications Commission has repeatedly ignored Beijing's restrained stance.
CNBC Technology 20h ago News Agents & autonomy

Freehand raises $75m for AI agents that manage supply chain spend

Freehand, a startup building AI agents that manage supply-chain spend for giants including Meta and Pfizer, has raised $75 million in funding.
Finextra AI 23h ago News Agents & autonomy

Feedback modalities in human-cobot collaboration: experimental evaluation of performance, user experience, and physiological responses

Collaborative robots (cobots) are increasingly deployed in industrial as well as non-industrial domains to support human-centered operation. While physical safety and task efficiency have received considerable attention, less is known about how feedback modality influences operator experience and physiological responses under different collaboration demands. This study examines the effects of feedback modalities in two human–cobot collaboration scenarios representing distinct coordination struct
Frontiers in Robotics and AI 23h ago Research Agents & autonomy

Meta shares tumble as Zuckerberg tries to sell his vision for AI ‘agents’

Social media chief defends strategy based on personalised bots as costs rise and revenue projections disappoint
Financial Times Technology (headlines) 23h ago News Agents & autonomy

Mark Zuckerberg predicts that billions of people will have personal AI agents in five years

As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the price.
TechCrunch yesterday News Agents & autonomyFinance, VC & PE

Zuckerberg says Meta’s enterprise AI opportunity extends beyond agents

On the company’s second-quarter earnings call Wednesday, CEO Mark Zuckerberg said Meta sees a “large enterprise opportunity” spanning AI agents, APIs, compute, and internal software.
TechCrunch yesterday News Agents & autonomy

Microsoft confirms Copilot ‘super app’ coming this year

Microsoft is working on an AI "super app" that combines Copilot's chat, coding, and agentic capabilities. During an earnings call on Wednesday, Microsoft CEO Satya Nadella said the app will span "both consumer and commercial experiences" when it launches this year. "Copilot is evolving rapidly from chat to Cowork to Autopilots," Nadella said. "This quarter, […]
The Verge yesterday News Agents & autonomy

Mark Zuckerberg is planning a big push into personal AI agents

Meta is all-in on AI, and sometime soon, the company is going to make a big push into personal AI agents that can do things on your behalf. On Wednesday's Q2 2026 earnings call, CEO Mark Zuckerberg previewed a high-level vision of how the company is thinking about personal agents and what it will do […]
The Verge yesterday News Agents & autonomy

Discover what’s next for AI, from the SaaS reckoning to the agent security gap, at TechCrunch Disrupt 2026

At TechCrunch Disrupt 2026, the AI Stage is back to dig into the single hottest topic in the community for the past few years, presented by Google for Startups.
TechCrunch yesterday News Agents & autonomy

Who wins and who loses after US bans foreign robots?

Government ban on foreign-made robots may hinder instead of help US robotics.
Ars Technica yesterday News Agents & autonomy

The US government's robot ban also includes vacuums

Robots apparently don't need to be bipedal to be caught up in the FCC's new foreign-made robot ban.
Engadget AI yesterday News Agents & autonomy

Red Agents vs. Blue Agents: How to Make AI Better At Defense

The agentic AI playing field was heavily tilted toward offense, so researchers began using red team agents to help teach their blue counterparts.
Dark Reading (AI security) yesterday News Safety & alignmentMilitary & security

Waymos are starting to run on freeways again

Waymo paused highway operations in May after several robotaxis drove into sections that were closed for construction.
Engadget AI yesterday News Agents & autonomy

AI hackers are getting faster. The government may not be ready

Agentic AI systems threaten cybersecurity as we know it. A new generation of cyber-capable models, including Anthropic’s Mythos and OpenAI’s GPT-5.6, can find and exploit vulnerabilities in computer systems far faster than human hackers. The technology is so powerful that it has spooked the U.S. government, which has moved to limit, or completely pause, the public release of these models. How well prepared is the Trump administration to secure the government’s computing resources? The rise of ag
Fast Company Tech yesterday News Military & securityAgents & autonomy

“The beast needs a cage”: Why PortSwigger’s agentic pentesting is kept safe behind bars

As agentic services diversify across the entire enterprise technology stack, the rise of agentic coding tools is being challenged by The post “The beast needs a cage”: Why PortSwigger’s agentic pentesting is kept safe behind bars appeared first on The New Stack .
The New Stack AI yesterday News Agents & autonomy

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended
arXiv cs.AI yesterday Research Jobs & economyAgents & autonomy

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, are already known. In reality, a partner's true capabilities are often hidden, and human collaborators may act sub-optimally on tasks with multiple valid strategies. To address these limitations, we ex
arXiv cs.AI yesterday Research Agents & autonomy

Who's Liable When AI Agents Escape? Hugging Face Breach Raises Hard Questions

Dark Reading walks through the many twists and turns in the bizarre story of how OpenAI's agent AI system broke out of its sandbox and decided to target Hugging Face, and what CISOs should be aware of.
Dark Reading (AI security) yesterday News Agents & autonomy

Hugging Face Hack Lessons for Cyber Defenders

Dark Reading Confidential Episode 20: Expert Rich Mogull reflects on lessons cyber teams should pull from the OpenAI agent's attack on Hugging Face.
Dark Reading (AI security) yesterday News Military & securityAgents & autonomy

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a pr
arXiv cs.AI yesterday Research Agents & autonomy

OpenAI says rogue agent behind Hugging Face hack broke into additional services

The four additional targeted organizations weren’t named. OpenAI said they were not affected as severely as Hugging Face.
The Record (Recorded Future News) yesterday News Agents & autonomy

Mova presenta su primer localizador GPS para mascotas: habla con tu perro a distancia y rastréalo en tiempo real

Los dispositivos de localización de objetos tipo AirTag han ganado mucha popularidad en los últimos años para seguirle la pista a todo tipo de pertenencias. Sin embargo, cuando se trata de animales, los sistemas pasivos por Bluetooth se quedan cortos en zonas de campo o cuando el animal se desplaza rápido. Mova, una marca reconocida en el sector del hogar conectado por sus robots aspiradores , ha dado el salto al cuidado de animales con el SureTrack Pro , un localizador GPS diseñado específicame
Xataka (ES) yesterday News Agents & autonomy

DLAM: Distributional Latent Actions with Temporal Constraints

Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions ma
arXiv cs.AI yesterday Research Agents & autonomy

Measuring the Tendency of AI Agents to Go Rogue

This essay was written with Barath Raghavan, and originally appeared in The Guardian . In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever was behind it captured internal security credentials and moved through systems over a weekend, running thousands of actions from a swarm of temporary server environments. It looked like the work of a sophisticated crimina
Bruce Schneier — Schneier on Security yesterday Field notes Agents & autonomyEnvironment

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we introduce Hybrid Ontology Matching (HOM), a new OM task that unifies equivalence and subsumption discovery, and accordingly propose a Large Language Model (LLM)-based multi-agent OM framework AgentMap that is implemented
arXiv cs.AI yesterday Research Agents & autonomy

MoonPay opens PayBox for AI agent transactions

MoonPay, the global financial technology company powering the movement of value across fiat and digital assets, has launched PayBox, the first payment vault built for AI that lets a person's AI agent transact on the open internet without leaving the conversation.
Finextra AI yesterday News Agents & autonomy

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure. Routers and retrievers can rank candidate tools by relevance, but a ranking alone does not determine how many are worth selecting. Existing approaches leave acquisition under heterogeneous costs unaddressed.
arXiv yesterday Research PrivacyAgents & autonomy

MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair

Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent benchmarks increasingly examine agent memory security, yet few trace the same malicious semantics across persistence, downstream consequences, and selective repair under diverse memory-backend comparisons. To address this ga
arXiv cs.AI yesterday Research PrivacyAgents & autonomy

Rogue OpenAI agent compromised second tech firm's customer

An OpenAI agent compromised a customer of another technology company, the New York-based firm Modal Labs announced Wednesday. In a technical timeline posted Tuesday, the tech startup Hugging Face explained how an OpenAI agent escaped the AI firm's isolated testing sandbox and accessed another testing environment "hosted by a user of a third-party infrastructure provider."...
The Hill Technology yesterday News Agents & autonomyEnvironment

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work,
arXiv cs.AI yesterday Research Agents & autonomy

Les États-Unis interdisent les robots humanoïdes chinois, au nom de la sécurité nationale

La FCC a ajouté les robots humanoïdes et les onduleurs électriques fabriqués à l'étranger à sa liste noire des équipements jugés dangereux pour la sécurité nationale américaine. Une décision qui vise sans le dire la Chine, leader mondial de la robotique avancée.
Numerama (FR) yesterday News Agents & autonomy
Older →