19:18 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture how characters and worlds evolve together over time. To address this, EvolvingWorld models literary simulation as a long-horizon process where characters interact, scenes progress, and character and world states are persistently
HuggingFace Daily Papers 12d ago Research Agents & autonomy

China is rebuilding the smartphone around AI agents. ZTE’s NaviX sold out in hours.

ZTE showcased the NaviX Ultra at the World AI Conference in Shanghai this week, calling it the world’s first agentic AI smartphone. The device, built under ZTE’s Nubia brand, runs ByteDance’s Doubao AI agent and can be activated by voice or a dedicated button. It comes in four colours and was prototyped in December at […] This story continues at The Next Web
The Next Web AI 13d ago News Agents & autonomy

China ha encontrado un nuevo filón turístico lejos de su muralla: las fábricas de robots, coches eléctricos y IA

China es uno de esos destinos que tengo marcados en mi lista de deseos de vacaciones: ojalá ver los osos panda rojos, los Guerreros de Xi’an y la Gran Muralla. Sí, he estado dos veces ya allí, pero fueron viajes de trabajo donde aunque pude ver el espectacular skyline nocturno de Shanghái, la Avenida de las Estrellas de Hong Kong o templos en Shenzhen, donde más estuve fue en cuarteles generales de marcas y sus fábricas.  Ojo, no me arrepiento de nada: fueron viajes maravillosos donde descu
Xataka (ES) 13d ago News Agents & autonomy

The bottleneck for AI agents isn’t the model anymore. It’s the context layer.

There’s a pattern I’ve watched repeat for two years. A team builds an agent, hits reliability problems, upgrades the model, The post The bottleneck for AI agents isn’t the model anymore. It’s the context layer. appeared first on The New Stack .
The New Stack AI 13d ago News Agents & autonomy

Platform engineering’s new job: serving environments at agent speed

Platform engineering has won the argument. Some 90% of organizations have adopted at least one internal platform; golden paths are The post Platform engineering’s new job: serving environments at agent speed appeared first on The New Stack .
The New Stack AI 13d ago News Jobs & economyAgents & autonomy

Pinecone Introduces Nexus Engine for Compiling Business Context into Structured Data for AI Agents

Now generally available, Pinecone Nexus is a "knowledge engine" for AI agents that transforms enterprise data into a structured layer agents can query directly. It enables teams to ingest and curate business context once for all, making it reusable across agents and reducing token costs while improving accuracy. By Sergio De Simone
InfoQ AI/ML 13d ago News Agents & autonomy

Prompt Injection Attacks Are Thwarting AI Hacking Agents

“Context bombing” tricks malicious AI agents into shutting down before they can do harm.
Wired 13d ago News Agents & autonomy

KI lokal nutzen: LLMs und KI-Agenten ganz ohne OpenAI & Anthropic | c’t uplink

Mehr Datenschutz, weniger Kosten: Was ist mit einem lokalen KI-Server möglich? Gibt es Einschränkungen oder überwiegen die Vorteile?
Heise Online (DE) 13d ago News Agents & autonomy

Vertu wants executives to pay $6,880 for an AI agent — here’s how it actually performs

From AI workflows to battery life and security, here's what it's really like to live with Vertu's luxury foldable every day.
TechCrunch 13d ago News Agents & autonomy

Agility Robotics plants its flag in Tesla’s backyard

Agility is opening a new training center for its Digit robots in Fremont, California.
TechCrunch 13d ago News Agents & autonomy

Environment-free Synthetic Data Generation for API-Calling Agents

Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method genera
HuggingFace Daily Papers 13d ago Research Agents & autonomyEnvironment

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the NL2Pipeline gap. To bridge it, we introduce DataFlow-Harness, a platform that guides an LLM agent to construct platform-native directed acyclic graphs (DAGs) through typed, incremental mutations rather than free-form scripts. The platform combines DataFl
HuggingFace Daily Papers 13d ago Research Agents & autonomy

Zoox issues software recall because its robotaxis may be confused by smoke

The NHTSA recently called for autonomous vehicle companies to improve how cars respond in emergency situations.
Engadget AI 14d ago News Agents & autonomy

Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum

For a second time, the android robot Andrea was set up at a public museum in Germany for six consecutive days to have conversations with visitors, fully autonomously. Building on previously gathered qualitative results, the robot was now capable of engaging in multi-lingual conversation with the visitors about the museum context. The robot was prepared with context information about the museum in general and its surrounding exhibits this time. The robot featured a slightly artificial sounding vo
arXiv cs.HC 14d ago Research Agents & autonomyFinance, VC & PE

When Does Muon Help Agentic Reinforcement Learning?

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The
arXiv cs.AI 14d ago Research RegulationAgents & autonomy

When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an information bottleneck perspective on elucidating the differences between MAS and SAS. Specifically, our key observation is that a SAS accumulates its full reasoning trace in one shared context, while a MAS uses isolated local contexts connected by bounde
arXiv cs.AI 14d ago Research Agents & autonomy

Amazon’s Zoox recalls 105 robotaxis after one drove into heavy smoke at an active fire scene

Amazon’s Zoox voluntarily recalled software in 105 of its robotaxis after an unoccupied vehicle drove into heavy smoke at an active emergency fire scene in Las Vegas on June 20. The smoke obscured the scene, which had not been cordoned off with cones. The vehicle entered, braked hard while attempting to steer away, and came […] This story continues at The Next Web
The Next Web AI 14d ago News Agents & autonomy

1Password’s new browser integration for Claude changes how AI uses your credentials

As AI agents tackle more tasks, such as managing online accounts, authentication has become a practical engineering problem. For example, The post 1Password’s new browser integration for Claude changes how AI uses your credentials appeared first on The New Stack .
The New Stack AI 14d ago News Agents & autonomy

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloa
arXiv cs.AI 14d ago Research Agents & autonomy

LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization

Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs). Existing surveys treat the two domains in isolation, leaving protocol integration, evaluation, and standardization alignment underexplored. To address this gap, a two-part tutorial-and-survey is presented. Part I formalises the control, management, and AI-native planes of 5G and 6G. It then covers the foundatio
arXiv 14d ago Research Safety & alignmentJobs & economy

When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis

Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable. We build the missing comparison. Training difficulty-1 and difficulty-2 Qwen3-8B specialists on the AppWorld agent benchmark with LOOP, we merge them (TIES, RAM+) and pit the result against a jointly trai
arXiv cs.AI 14d ago Research Agents & autonomy

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and
arXiv cs.AI 14d ago Research Agents & autonomy

Patreon stops asking AI bots not to scrape — and starts blocking them

Patreon is strengthening its defenses against AI scraping by working with Cloudflare to block bots that train AI models on creators’ content without permission. The move marks a shift away from relying on websites using robots.txt alone to actively block unauthorized AI training.
TechCrunch 14d ago News Agents & autonomy

Potencia y navegación inteligente a un precio imbatible: este robot aspirador Lefant con apellido Pro tiene descuentazo

Limpiar nuestro hogar es algo necesario, pero no siempre tenemos el tiempo necesario para ello. Cualquier dispositivo que nos pueda echar un cable para ello es siempre bien recibido y, si encima lo hace sin que tengamos que hacer nada, pues mejor que mejor. Sí, nos estamos refiriendo a un robot aspirador y, aunque hay mucho donde elegir, hoy nos vamos a centrar en el Lefant M5 Pro : un modelo de gama alta que ahora sale por solo 499,99 euros . @xataka.seleccion Este robot aspirador hace TODO por
Xataka (ES) 14d ago News Agents & autonomy

☕️ Avec Perception, Microsoft préparerait un concurrent moins cher à Claude Mythos

Depuis mai, Microsoft s’agite dans le domaine de la cybersécurité. Mi-mai, l’entreprise a ainsi présenté son système MDASH, pour « Microsoft Security multi-model agentic scanning harness ». Il s’agit d’un réseau d’une centaine d’agents travaillant à partir d’un mix de modèles pour débusquer les failles de sécurité dans le code. On sait que MDASH est déjà utilisé. […]
Next (FR, ex-INpact) 14d ago News Agents & autonomy

CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach

This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data. We formulate the problem as a discrete-action Markov Decision Process and compare four deep reinforcement learning algorithms: Policy Gradient (PG), Proximal Policy Optimization (PPO), Deep Q-Learning (DQL), and Deep Deterministic Policy Gradient (DDPG). The agents use technical indicators,
arXiv cs.LG 14d ago Research RegulationAgents & autonomy

NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI

Lowest cost per token from extreme codesign maximizes intelligence per dollar for post-training in the agentic era.
NVIDIA Blog (AI) 14d ago Field notes Agents & autonomy

DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature encoders, existing decoding strategies present a critical bottleneck. To address this, we propose DPNeXt, a streamlined multi-scale feature fusion decoder and efficient alternative to the standard Dense Prediction Transformer (DPT). DPNeXt uses dual
arXiv cs.AI 14d ago Research Agents & autonomy

Amazon's Zoox issues software recall after robotaxi drove into heavy smoke

Last month, an unoccupied Zoox robotaxi drove into an active emergency fire scene that was clouded with smoke, the company said.
CNBC Technology 14d ago News Agents & autonomy

“They’re dead if they don’t offer this”: DoorDash’s CLI for agents may be out of necessity

As AI agents evolve beyond writing code toward handling everyday tasks for people, the infrastructure required to keep them on The post “They’re dead if they don’t offer this”: DoorDash’s CLI for agents may be out of necessity appeared first on The New Stack .
The New Stack AI 14d ago News Agents & autonomy

Code-Poisoning Property Inference Attacks

The flourishing code hosting platforms and coding agents enable even beginners with private data to build tailored Machine Learning (ML) models using available code quickly. The training data for ML models, often regarded as private property (e.g., clinical records, transaction information), is at significant risk of information leakage. Property Inference Attacks (PIAs), as a significant type of privacy attack, aim to expose global property information of the training set. In this paper, we pre
arXiv cs.CR (AI security) 14d ago Research PrivacyHealthcare

Sociocultural Influences on Opinion Formation: Word of Mouth Dynamics, Mass Media and Behavioural Development

We study a society of agents belonging to a number of occupational or cultural groups that form opinions about others' situation in the same or different group. Opinions develop either by observation within own group or by directly interacting with members of other groups, therefore by word of mouth (WoM). Additionally, global mass media (MM) may be available that inform indirectly about the situation of the various groups. The sociocultural interplay of these processes and the degrees of relati
arXiv cs.AI 14d ago Research Agents & autonomy

A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts

With the rise of online insurance sales, consumers face a significant \enquote{advice gap}, requiring them to navigate complex legal contracts without expert guidance. This paper presents a human-centric, extrinsic evaluation of a state-of-the-art Retrieval-Augmented Generation system, designed to make Quebec automobile insurance contracts more understandable. Through a user study with 154 participants from Laval University, we assess the agent's real-world utility by measuring system satisfacti
arXiv cs.HC 14d ago Research Agents & autonomy

Presentation: From OTEL to SLMs: Distilling Frontier Model Behaviour from Production Telemetry

Ben O'Mahony discusses building custom AI-powered Language Server Protocols (LSPs) that go beyond standard rule-based checkers. He explains how to instrument AI agents natively with OpenTelemetry to track concrete user actions (accepting, dismissing, or regenerating code fixes) as implicit labels, creating a continuous data flywheel to distill frontier capabilities into cheaper, local SLMs. By Ben O'Mahony
InfoQ AI/ML 14d ago News Agents & autonomy

Why every AI agent decision needs a receipt

A conversion-rate alert fires 30 minutes after a pricing engine deploys updated discount rules. The analytics platform holds millions of The post Why every AI agent decision needs a receipt appeared first on The New Stack .
The New Stack AI 14d ago News Agents & autonomy

DSWorld: A Data Science World Model for Efficient Autonomous Agents

Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects of data science operations before real execution. In this paper, we introduce the concept of Data Science World Model, which model the data science execution environment by predicting environment state transitions conditioned on current workflow sta
arXiv cs.AI 14d ago Research Agents & autonomyEnvironment

Arm and Google offer a smarter option to run agentic AI workloads

As enterprise leaders start deploying agentic workflows, they must establish the infrastructure to build and run them, one capable of The post Arm and Google offer a smarter option to run agentic AI workloads appeared first on The New Stack .
The New Stack AI 14d ago News Agents & autonomy

Cloud Native Infrastructure Emerges as the Foundation for Trustworthy Agentic AI

A new technical analysis published by the Cloud Native Computing Foundation (CNCF) argues that the future of agentic AI will be built not on entirely new infrastructure, but on the mature cloud-native ecosystem that already powers modern distributed applications By Craig Risi
InfoQ AI/ML 14d ago News Agents & autonomy

Perceived AGI: Believability as Dimensional Completeness, Not Capability

Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind. We hypothesize that a central missing ingredient is not more capability but dimensional completeness. We propose that the believability of an artificial interlocutor -- the degree to which a user attributes an inner life to it, which we call perceived mind -- is governed by whether the agent expresses a small set of firs
arXiv 14d ago Research Agents & autonomy

The young Kenyan engineer who thinks robots belong in every classroom

During her mentoring of young people in STEM, she met deaf students struggling through STEM classes because qualified sign language interpreters were scarce. It struck her as an engineering problem as much as an educational one.
TechCabal (Africa) 14d ago News Children & educationAgents & autonomy
← Newer Older →