Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture how characters and worlds evolve together over time. To address this, EvolvingWorld models literary simulation as a long-horizon process where characters interact, scenes progress, and character and world states are persistently
China is rebuilding the smartphone around AI agents. ZTE’s NaviX sold out in hours.
ZTE showcased the NaviX Ultra at the World AI Conference in Shanghai this week, calling it the world’s first agentic AI smartphone. The device, built under ZTE’s Nubia brand, runs ByteDance’s Doubao AI agent and can be activated by voice or a dedicated button. It comes in four colours and was prototyped in December at […] This story continues at The Next Web
China ha encontrado un nuevo filón turístico lejos de su muralla: las fábricas de robots, coches eléctricos y IA
China es uno de esos destinos que tengo marcados en mi lista de deseos de vacaciones: ojalá ver los osos panda rojos, los Guerreros de Xi’an y la Gran Muralla. Sí, he estado dos veces ya allí, pero fueron viajes de trabajo donde aunque pude ver el espectacular skyline nocturno de Shanghái, la Avenida de las Estrellas de Hong Kong o templos en Shenzhen, donde más estuve fue en cuarteles generales de marcas y sus fábricas. Ojo, no me arrepiento de nada: fueron viajes maravillosos donde descu
The bottleneck for AI agents isn’t the model anymore. It’s the context layer.
There’s a pattern I’ve watched repeat for two years. A team builds an agent, hits reliability problems, upgrades the model, The post The bottleneck for AI agents isn’t the model anymore. It’s the context layer. appeared first on The New Stack .
Platform engineering’s new job: serving environments at agent speed
Platform engineering has won the argument. Some 90% of organizations have adopted at least one internal platform; golden paths are The post Platform engineering’s new job: serving environments at agent speed appeared first on The New Stack .
Pinecone Introduces Nexus Engine for Compiling Business Context into Structured Data for AI Agents
Now generally available, Pinecone Nexus is a "knowledge engine" for AI agents that transforms enterprise data into a structured layer agents can query directly. It enables teams to ingest and curate business context once for all, making it reusable across agents and reducing token costs while improving accuracy. By Sergio De Simone
Prompt Injection Attacks Are Thwarting AI Hacking Agents
“Context bombing” tricks malicious AI agents into shutting down before they can do harm.
KI lokal nutzen: LLMs und KI-Agenten ganz ohne OpenAI & Anthropic | c’t uplink
Mehr Datenschutz, weniger Kosten: Was ist mit einem lokalen KI-Server möglich? Gibt es Einschränkungen oder überwiegen die Vorteile?
Vertu wants executives to pay $6,880 for an AI agent — here’s how it actually performs
From AI workflows to battery life and security, here's what it's really like to live with Vertu's luxury foldable every day.
Agility Robotics plants its flag in Tesla’s backyard
Agility is opening a new training center for its Digit robots in Fremont, California.
Environment-free Synthetic Data Generation for API-Calling Agents
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method genera
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the NL2Pipeline gap. To bridge it, we introduce DataFlow-Harness, a platform that guides an LLM agent to construct platform-native directed acyclic graphs (DAGs) through typed, incremental mutations rather than free-form scripts. The platform combines DataFl
Zoox issues software recall because its robotaxis may be confused by smoke
The NHTSA recently called for autonomous vehicle companies to improve how cars respond in emergency situations.
Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum
For a second time, the android robot Andrea was set up at a public museum in Germany for six consecutive days to have conversations with visitors, fully autonomously. Building on previously gathered qualitative results, the robot was now capable of engaging in multi-lingual conversation with the visitors about the museum context. The robot was prepared with context information about the museum in general and its surrounding exhibits this time. The robot featured a slightly artificial sounding vo
When Does Muon Help Agentic Reinforcement Learning?
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The
When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an information bottleneck perspective on elucidating the differences between MAS and SAS. Specifically, our key observation is that a SAS accumulates its full reasoning trace in one shared context, while a MAS uses isolated local contexts connected by bounde
Amazon’s Zoox recalls 105 robotaxis after one drove into heavy smoke at an active fire scene
Amazon’s Zoox voluntarily recalled software in 105 of its robotaxis after an unoccupied vehicle drove into heavy smoke at an active emergency fire scene in Las Vegas on June 20. The smoke obscured the scene, which had not been cordoned off with cones. The vehicle entered, braked hard while attempting to steer away, and came […] This story continues at The Next Web
1Password’s new browser integration for Claude changes how AI uses your credentials
As AI agents tackle more tasks, such as managing online accounts, authentication has become a practical engineering problem. For example, The post 1Password’s new browser integration for Claude changes how AI uses your credentials appeared first on The New Stack .
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloa
LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization
Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs). Existing surveys treat the two domains in isolation, leaving protocol integration, evaluation, and standardization alignment underexplored. To address this gap, a two-part tutorial-and-survey is presented. Part I formalises the control, management, and AI-native planes of 5G and 6G. It then covers the foundatio
When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable. We build the missing comparison. Training difficulty-1 and difficulty-2 Qwen3-8B specialists on the AppWorld agent benchmark with LOOP, we merge them (TIES, RAM+) and pit the result against a jointly trai
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and
Patreon stops asking AI bots not to scrape — and starts blocking them
Patreon is strengthening its defenses against AI scraping by working with Cloudflare to block bots that train AI models on creators’ content without permission. The move marks a shift away from relying on websites using robots.txt alone to actively block unauthorized AI training.
Potencia y navegación inteligente a un precio imbatible: este robot aspirador Lefant con apellido Pro tiene descuentazo
Limpiar nuestro hogar es algo necesario, pero no siempre tenemos el tiempo necesario para ello. Cualquier dispositivo que nos pueda echar un cable para ello es siempre bien recibido y, si encima lo hace sin que tengamos que hacer nada, pues mejor que mejor. Sí, nos estamos refiriendo a un robot aspirador y, aunque hay mucho donde elegir, hoy nos vamos a centrar en el Lefant M5 Pro : un modelo de gama alta que ahora sale por solo 499,99 euros . @xataka.seleccion Este robot aspirador hace TODO por
☕️ Avec Perception, Microsoft préparerait un concurrent moins cher à Claude Mythos
Depuis mai, Microsoft s’agite dans le domaine de la cybersécurité. Mi-mai, l’entreprise a ainsi présenté son système MDASH, pour « Microsoft Security multi-model agentic scanning harness ». Il s’agit d’un réseau d’une centaine d’agents travaillant à partir d’un mix de modèles pour débusquer les failles de sécurité dans le code. On sait que MDASH est déjà utilisé. […]
CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach
This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data. We formulate the problem as a discrete-action Markov Decision Process and compare four deep reinforcement learning algorithms: Policy Gradient (PG), Proximal Policy Optimization (PPO), Deep Q-Learning (DQL), and Deep Deterministic Policy Gradient (DDPG). The agents use technical indicators,
NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI
Lowest cost per token from extreme codesign maximizes intelligence per dollar for post-training in the agentic era.
DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature encoders, existing decoding strategies present a critical bottleneck. To address this, we propose DPNeXt, a streamlined multi-scale feature fusion decoder and efficient alternative to the standard Dense Prediction Transformer (DPT). DPNeXt uses dual
Amazon's Zoox issues software recall after robotaxi drove into heavy smoke
Last month, an unoccupied Zoox robotaxi drove into an active emergency fire scene that was clouded with smoke, the company said.
“They’re dead if they don’t offer this”: DoorDash’s CLI for agents may be out of necessity
As AI agents evolve beyond writing code toward handling everyday tasks for people, the infrastructure required to keep them on The post “They’re dead if they don’t offer this”: DoorDash’s CLI for agents may be out of necessity appeared first on The New Stack .
Code-Poisoning Property Inference Attacks
The flourishing code hosting platforms and coding agents enable even beginners with private data to build tailored Machine Learning (ML) models using available code quickly. The training data for ML models, often regarded as private property (e.g., clinical records, transaction information), is at significant risk of information leakage. Property Inference Attacks (PIAs), as a significant type of privacy attack, aim to expose global property information of the training set. In this paper, we pre
Sociocultural Influences on Opinion Formation: Word of Mouth Dynamics, Mass Media and Behavioural Development
We study a society of agents belonging to a number of occupational or cultural groups that form opinions about others' situation in the same or different group. Opinions develop either by observation within own group or by directly interacting with members of other groups, therefore by word of mouth (WoM). Additionally, global mass media (MM) may be available that inform indirectly about the situation of the various groups. The sociocultural interplay of these processes and the degrees of relati
A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts
With the rise of online insurance sales, consumers face a significant \enquote{advice gap}, requiring them to navigate complex legal contracts without expert guidance. This paper presents a human-centric, extrinsic evaluation of a state-of-the-art Retrieval-Augmented Generation system, designed to make Quebec automobile insurance contracts more understandable. Through a user study with 154 participants from Laval University, we assess the agent's real-world utility by measuring system satisfacti
Presentation: From OTEL to SLMs: Distilling Frontier Model Behaviour from Production Telemetry
Ben O'Mahony discusses building custom AI-powered Language Server Protocols (LSPs) that go beyond standard rule-based checkers. He explains how to instrument AI agents natively with OpenTelemetry to track concrete user actions (accepting, dismissing, or regenerating code fixes) as implicit labels, creating a continuous data flywheel to distill frontier capabilities into cheaper, local SLMs. By Ben O'Mahony
Why every AI agent decision needs a receipt
A conversion-rate alert fires 30 minutes after a pricing engine deploys updated discount rules. The analytics platform holds millions of The post Why every AI agent decision needs a receipt appeared first on The New Stack .
DSWorld: A Data Science World Model for Efficient Autonomous Agents
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects of data science operations before real execution. In this paper, we introduce the concept of Data Science World Model, which model the data science execution environment by predicting environment state transitions conditioned on current workflow sta
Arm and Google offer a smarter option to run agentic AI workloads
As enterprise leaders start deploying agentic workflows, they must establish the infrastructure to build and run them, one capable of The post Arm and Google offer a smarter option to run agentic AI workloads appeared first on The New Stack .
Cloud Native Infrastructure Emerges as the Foundation for Trustworthy Agentic AI
A new technical analysis published by the Cloud Native Computing Foundation (CNCF) argues that the future of agentic AI will be built not on entirely new infrastructure, but on the mature cloud-native ecosystem that already powers modern distributed applications By Craig Risi
Perceived AGI: Believability as Dimensional Completeness, Not Capability
Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind. We hypothesize that a central missing ingredient is not more capability but dimensional completeness. We propose that the believability of an artificial interlocutor -- the degree to which a user attributes an inner life to it, which we call perceived mind -- is governed by whether the agent expresses a small set of firs
The young Kenyan engineer who thinks robots belong in every classroom
During her mentoring of young people in STEM, she met deaf students struggling through STEM classes because qualified sign language interpreters were scarce. It struck her as an engineering problem as much as an educational one.