23:19 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial videos and other multimodal human resources largely underused. We present RESOURCE2SKILL, a framework that distills multimodal resources, including tutorial videos, repositories, articles, and reference artifacts, into executable skills for software agents. RES
HuggingFace Daily Papers 16d ago Research Agents & autonomy

Multi-Turn On-Policy Distillation with Prefix Replay

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed pr
HuggingFace Daily Papers 16d ago Research RegulationChildren & education

Amid hardware legal battle, OpenAI releases a $230 keyboard for Codex

OpenAI, which is in the middle of a legal battle with Apple over hardware trade theft allegations, just released a light-up keyboard designed to be paired with its agentic coding app.
TechCrunch 16d ago News Agents & autonomy

Traccia: An OpenTelemetry-Based Governance Platform for AI Systems

The rapid development of Large Language Models (LLMs) and Artificial Intelligent (AI) powered autonomous agents has fundamentally changed the existing forms of software governance. In spite of the rigorous standards of transparency and account ability required according to the international frameworks such as the European Union's AI Act, there is a considerable gap between theory and reality. The present study discusses the inherent drawbacks of currently utilized platforms for LLM evaluation, m
arXiv 16d ago Research RegulationAgents & autonomy

ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs

Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict? We investigate this question in the context of tool-calling LLM agents deployed in regulated industries, where agents processing confidential documents may encounter content that triggers safety-trained values (e.g., public welfare) that conflict with deployment-context instructions (e.g., internal logging). To empirically verify this phenomenon, we build a benchmark of 128 scena
arXiv 16d ago Research RegulationSafety & alignment

Behavior Change Content and Implementation of Large Language Model–Driven Conversational Agents in Cardiometabolic Care: Scoping Review

Background: Large language models (LLMs) are increasingly embedded in conversational agents for cardiometabolic care. These systems could support self-management, but their behavior change content, delivery mechanisms, and implementation transparency are poorly understood. Objective: This scoping review mapped behavior change techniques (BCTs) used in LLM-driven conversational agents for cardiometabolic prevention and management, described how these techniques are delivered across static, rule-b
JMIR (Journal of Medical Internet Research) 16d ago Research Agents & autonomyTransparency

AI Agents Do Not Fail Alone:The Context Fails First

Context engineering has become central to building reliable AI agents, yet it remains largely unmeasured. Agents do not fail in isolation: their behavior is shaped by the instructions, tools, memory, retrieved knowledge, guardrails, and untrusted inputs accumulated in their context. When this context is weak, agents drift, hallucinate, misuse tools, ignore constraints, become vulnerable to injection, and waste tokens. This paper validates context-engineering quality as an independent leading ind
arXiv 16d ago Research Agents & autonomy

OpenAI launches a physical keypad for controlling agents

OpenAI's collaboration with keyboard maker Work Louder is available to order today.
Engadget AI 16d ago News Agents & autonomy

Kubernetes won the container decade. Google’s Agent Substrate wants the next one.

Google made GKE Agent Sandbox generally available in May 2026 and, in the same post, introduced a second project called The post Kubernetes won the container decade. Google’s Agent Substrate wants the next one. appeared first on The New Stack .
The New Stack AI 16d ago News Agents & autonomy

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation

Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases. Traditional active learning and manual annotation fail to scale against the complexity and volume of novel multimodal threats. In this paper, we propose an automated, agentic red-teaming framework that systematically synthesizes difficult examples using an iterative strategy that proposes novel hy
arXiv red teaming query 16d ago Research Safety & alignmentAgents & autonomy

Trust, transactions and tokenomics: AI agent infrastructure begins to standardize

As AI agents gain greater autonomy across the internet, a system of governance is emerging around three questions: how much The post Trust, transactions and tokenomics: AI agent infrastructure begins to standardize appeared first on The New Stack .
The New Stack AI 16d ago News RegulationAgents & autonomy

« Je n’ai plus qu’à lui dire adieu » : la Chine veut juguler les IA simulant les relations humaines

Les internautes chinois ont massivement adopté les agents IA pour des relations amicales ou amoureuses. Une tendance contre laquelle Pékin tente désormais de lutter en imposant, mercredi, de nouvelles restrictions.
Le Monde Pixels (FR) 16d ago News Agents & autonomy

Air Force reaches CCA milestone with live-firing of missile from Anduril’s robotic fighter jet

“We’re one step closer to delivering capabilities to the warfighter,” Air Force Chief of Staff Gen. Ken Wilsbach said in a statement. The post Air Force reaches CCA milestone with live-firing of missile from Anduril’s robotic fighter jet appeared first on DefenseScoop .
DefenseScoop 16d ago News Agents & autonomy

What building Shippy taught us about building agents

Hugging Face Blog 16d ago Field notes Agents & autonomy

ICE Reverses Plan to Halt Vehicle Stops After Trump Complains

The walk back came a day after ICE agents had been told to avoid vehicle stops in the wake of deadly shootings in Maine and Houston.
Time Tech 16d ago News Agents & autonomy

PhysClaw-0: A Symbiotic Agentic System for Robot Autonomy via Language Corrections

Autonomous data collection governs the volume and quality of real-world trajectories for manipulation policy learning. Existing pipelines reduce human effort via self-resetting, VLM verification, or language-guided correction, yet episode-scoped fixes must be reissued whenever the same failure recurs, so oversight cost grows with session length rather than with the number of distinct problems. We present PhysClaw-0, a human-robot symbiotic agentic system in which corrections are retained and reu
arXiv cs.HC 16d ago Research RegulationAgents & autonomy

Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment

Autonomous data collection governs the volume and quality of real-world trajectories for manipulation policy learning. Existing pipelines reduce human effort via self-resetting, VLM verification, or language-guided correction, yet episode-scoped fixes must be reissued whenever the same failure recurs, so oversight cost grows with session length rather than with the number of distinct problems. We present Zero2Skill, a human-robot symbiotic agentic system in which corrections are retained and reu
arXiv cs.HC 16d ago Research RegulationAgents & autonomy

Crear un agente de IA que te ayude a automatizar procesos ya no exige saber programar: así es Hermes Agent de Hostinger

La IA ya forma parte de nuestra vida, tanto personal como laboral. En este último ámbito, puede ser una herramienta que nos ayude mucho con procesos de todo tipo, pero todo va a depender de qué utilicemos para ello. Hay muchos autónomos o emprendedores que simplemente buscan algo sencillo e intuitivo que no implique muchas complicaciones. Eso es justo lo que ofrece Hermes Agent de Hostinger, ahora con un ofertón : de 18,99 euros al mes pasa a costar 4,99 euros al mes. Una IA para automatizar y f
Xataka (ES) 16d ago News Agents & autonomy

Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education

This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a conversational AI assistant based on Retrieval-Augmented Generation. It aims to enhance earthquake preparedness and conscious action among primary-school students. The system extends the award-winning STEM project Earthquaker moving from mechanical simulation with Lego WeDo2 to cognitive and metacognitive processing. The robotics component uses L
arXiv 16d ago Research Children & educationAgents & autonomy

Early Adoption of Agentic Coding Tools by GitHub Projects

Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level outcomes of agent-generated contributions, less is known about how agentic coding tools are adopted and managed at the project level. In this paper, we analyze 25,264 agentic PRs from 2,361 popular GitHub repositories to investigate (1) the adoption of agentic cod
arXiv 16d ago Research Agents & autonomyFinance, VC & PE

Early Adoption of Agentic Coding Tools by GitHub Projects

Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level outcomes of agent-generated contributions, less is known about how agentic coding tools are adopted and managed at the project level. In this paper, we analyze 25,264 agentic PRs from 2,361 popular GitHub repositories to investigate (1) the adoption of agentic cod
arXiv 16d ago Research Agents & autonomyFinance, VC & PE

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new failures and new tasks appear over time. The central question this raises is whether optimizer-driven gains compound: after an agent has been optimized once, can it be optimized again
arXiv cs.AI 16d ago Research Agents & autonomy

The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce

The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms. As AI evolves from passive recommendation algorithms to autonomous, goal-directed agents capable of executing purchasing decisions, the conventional understanding of consumer-brand relationships requires a structural reevaluation. By synthesizing extant literature across human-machine teaming, consumer decision-making, and algorithmic trust dynamics, we demonstrate that tradi
arXiv 16d ago Research Agents & autonomy

Atlassian evolves Jira into an orchestration hub for developers and AI agents

Atlassian Corp. today announced it is expanding Jira with updates that will help developers prepare, distribute and track work performed by artificial intelligence agents. The company’s new Jira Planner helps turn incomplete project ideas into technical specifications, while its Jira Coding Agent and integrations with third-party agents transform work items into requests. With automation rules […] The post Atlassian evolves Jira into an orchestration hub for developers and AI agents appeared fir
SiliconANGLE AI 16d ago News Jobs & economyAgents & autonomy

OpenAI's first branded hardware is... a light-up keyboard?

The Codex Micro is designed to monitor multiple agentic threads at a glance.
Ars Technica 16d ago News Agents & autonomy

Perplexity launches secure sandbox to make its AI agents secure and powerful

Perplexity AI Inc. today introduced a new feature that takes its current agentic artificial intelligence service, Computer, to perform better with greater security. The company introduced SPACE, a sandbox platform designed to allow its AI agent to act with its full capabilities, while providing the highest level of security for agentic systems. Perplexity Computer can […] The post Perplexity launches secure sandbox to make its AI agents secure and powerful appeared first on SiliconANGLE .
SiliconANGLE AI 16d ago News Agents & autonomy

Android Studio Quail 2 mit parallelen Agentenchats und LeakCanary-Integration

Googles offizielle IDE für die Entwicklung von Android-Apps setzt verstärkt auf KI-Integration, unter anderem mit mehreren parallel laufenden Agentenchats.
Heise Online (DE) 16d ago News Agents & autonomy

Claude Flaw Automatically Sends Malicious Prompts to AI Agents

When combined with another exploit, the "PromptFiction" vulnerability, which has been fixed, could have enabled an end-to-end attack on a targeted system.
Dark Reading (AI security) 16d ago News Agents & autonomy

A Self-Evolving Agent for Longitudinal Personal Health Management

Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an open-source agent architecture that updates support as a person's routines, preferences, measurements and risks change. It separates shared safety rules and medical knowledge from private longitudinal memory containing profile facts, reusable procedures and episodic traces. After each episode, induction determines what should update the profile, rev
arXiv 16d ago Research HealthcareAgents & autonomy

ExpressionCueLens: A Cross-Cultural Analysis of Human-AI Companion Conversations on Social Media

LLM-based AI companion agents are increasingly being perceived not only as tools but also as social companions. On social media, people recount conversations where these agents comfort, negotiate and assert boundaries, reflecting a growing attribution of human-like qualities. To profile how agency is perceived in human-AI (HAI) interactions, we introduce the ExpressionCueLens framework, which organizes linguistic, cognitive, behavioral and perceptual cues into ten categories of anthropomorphism
arXiv cs.HC 16d ago Research Agents & autonomy

Experience Memory Graph: One-Shot Error Correction for Agents

Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of states, actions, and observations. However, in complex, long-horizon tasks, these agents frequently suffer from compounding errors and struggle to recover from failures. Existing self-correction mechanisms rely on prompt-based reflection, which is inherently brittle, incurs heavy time and API costs due to iterative trial-and-error loops, and produces task-sp
arXiv cs.AI 16d ago Research Agents & autonomy

Stripe Benchmark Shows AI Agents Build Integrations but Struggle with Validation

Stripe introduces a benchmark suite to evaluate whether AI agents can build real-world Stripe integrations across backend, frontend, and browser-based checkout workflows. The study examines end-to-end software engineering capability, focusing on execution, testing, and validation gaps in agentic systems under production-like constraints. By Leela Kumili
InfoQ AI/ML 16d ago News Agents & autonomy

Paytia lets AI agents take card payments — without the card ever reaching the AI

Businesses are putting AI agents on the phone and in chat faster than their compliance teams can keep up.
Finextra AI 16d ago News RegulationAgents & autonomy

Anaconda buys Kilo, the open source coding agent that answers to no single model maker

Anaconda, a company that provides governed, open-source packages and environments for enterprises, has acquired popular open-source coding agent Kilo. The The post Anaconda buys Kilo, the open source coding agent that answers to no single model maker appeared first on The New Stack .
The New Stack AI 16d ago News Agents & autonomyEnvironment

How Bolna AI Is Helping Enterprises Win The Voice AI Race

Enterprises are moving fast to adopt voice AI. Across ecommerce, banking, education, recruitment and customer support, companies want voice agents…
Inc42 (IN) 16d ago News Children & educationAgents & autonomy

Presentation: Postgres for Production Agents: Your Relational Foundation for Enterprise AI

Gwen Shapira shares how teams are scaling AI features using PostgreSQL for mission-critical apps. She explains how to leverage Postgres's multi-modal capabilities - including JSONB parsing and high-recall HNSW vector indexing - to deliver deterministic and semantic context to LLMs. She also discusses vector quantization to speed up queries by 4x and strategies for managing agentic memory. By Gwen Shapira
InfoQ AI/ML 16d ago News Agents & autonomy

Conjecture Machines

AI agents and the new validation bottleneck in science
AI Policy Perspectives 16d ago Field notes Agents & autonomy

Creatio expands beyond its no-code roots with conversational development tool and AI studio

Creatio Inc. today introduced what it calls a major update to its customer relationship management and workflow platform that enables business users and information technology teams to build, deploy and govern artificial intelligence agents alongside conventional customer workflows. The Creatio 10x release combines the company’s CRM applications with new tools for developing personal assistants, deterministic […] The post Creatio expands beyond its no-code roots with conversational development t
SiliconANGLE AI 16d ago News Agents & autonomy

Litera ‘Relaunches’ With One Agent to Rule the Platform

Legal tech legend, Litera, has ‘relaunched’ itself with its Lito agent as the principal point of focus, offering the ability to tap both its business ...
Artificial Lawyer 16d ago News Agents & autonomy

Vint Cerf is working on a plan to unleash AI agents on the open internet

The guy behind TCP/IP is working on a standard for identifying AI agents in the wild.
TechCrunch 16d ago News Agents & autonomy
← Newer Older →