Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
"Es un problema muy complejo de resolver": Tesla se desploma en bolsa, pero la culpa no es de los coches, sino de los robots
Las acciones de Tesla se han hundido este jueves entre un 12% y un 14% , siendo uno de sus mayores desplomes en los últimos años. La caída llegó tras la publicación de los resultados del segundo trimestre, agravándose incluso más durante la conferencia con inversores, en la que Elon Musk reconocía que fabricar sus robots humanoides Optimus a gran escala será mucho más complicado de lo que había prometido hasta ahora. Según datos recogidos
TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI
Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic applications execute as long-horizon workflows whose quality is determined only by a delayed, task-level outcome. This mismatch prevents per-call routers from correctly attributing feedback to individual routing decisions. Towards mitigating this, we pr
Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture
Enterprise AI agents are typically granted static credential sets at configuration time, holding every tool the role might need for every task they perform. This persistent over-privilege expands the attack surface. We argue that capability scoping must follow a dynamic least-privilege principle and be treated as a prevention mechanism before a detection one. A credential that does not exist in an agent's context cannot be misused regardless of the agent's reasoning or evasion sophistication. We
Pentagon eyes November to demo ground-launched, precision strike weapons
Aligned with the new autonomy czar shop, the department plans to spend $250 million evaluating options and moving out with initial deals for its new Ground-Based Affordable Mass initiative.
AWS, Google Cloud, Microsoft Azure, and Cloudflare now all offer agent sandboxes. None built them the same way.
Google Cloud announced it had put Cloud Run sandboxes into public preview earlier this month at WeAreDevelopers World Congress in The post AWS, Google Cloud, Microsoft Azure, and Cloudflare now all offer agent sandboxes. None built them the same way. appeared first on The New Stack .
Robot Learning to Communicate through Projected Visual Abstractions
Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through such projected visual abstractions requires reasoning not only about bodily motion but also about how that motion is transformed into an external representation perceived by an observer. Among these abstractions, shadows provide a particularly compel
A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation
Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from static conversational interfaces to dynamic systems capable of complex reasoning, tool execution, and decision-making. However, the operational reliability of these agentic AI systems is fundamentally challenged by the absence of reliable ground truth in open-ended environments and the risk of increasing operational drift over time. To address this challenge, we propose and experimentally evaluate an
SceneActBench: Can Agents Act on the 3D Scenes They See?
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D envi
Agentic Root Cause Analysis through Evidence-Grounded Reasoning
Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck. While existing data-driven approaches aim to automate this, two critical limitations restrict their deployment: their operate as black boxes unable to justify their diagnosis, and they require scarce labeled examples of faulty operation. To address this gap, we
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. This often leads to the generation of ideas in close proximity to one another or to a large set of trivial, unsound, or unclear concepts. In this work, we instead argue that research ideation should be treated as a conjunction of both objectives an
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a
The 3 types of people who will excel in the AI agent era, according to tech leaders
The autonomous business of the future is being built. Certain skills are in high demand - and can help you stand out.
Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education
Generative social robots (GSRs) powered by large language models offer new possibilities for personalized tutoring in higher education, but also introduce risks related to misinformation, missing transparency, or reinforcing incorrect student responses. Prior work identified knowledge-based design (KBD) requirements that define the informational prerequisites for GSRs to manifest responsible and effective tutoring behavior in higher education. In this paper, we operationalized selected KBD requi
Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG
Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and cost-efficient integration through knowledge-grounded LLMs and agents operating within a retrieval-augmented generation (RAG) workflow. Here, trustworthiness refers to evidence-ground
OpenAI’s new voice mode makes it to the ChatGPT desktop app
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
Un robot géant avec une cape : le prochain anime Gundam a déjà tout compris
La saga Gundam s'apprête à ouvrir un nouveau chapitre en 2027. Et, surprise de la taille d'un mécha : l'anime sera directement lié à un nouveau jeu vidéo Gundam, attendu lui aussi pour 2027.
Be skeptical of OpenAI’s rogue hacker agent story | John Thickstun
If OpenAI loudly proclaims how dangerous AI is, investors will hear how powerful it is. And who benefits from that? On 14 February 2019, OpenAI announced a language model called GPT-2, the precursor to the models that power modern AI chatbots and agents such as ChatGPT and Claude. But OpenAI declared GPT-2 was too risky to release, citing concerns about safety and abuse. I recall being annoyed at the time that OpenAI would make such a useless announcement: the risks seemed overblown, and without
When is an apology not an apology? When it comes from an AI boss with an out-of-control chatbot | Marina Hyde
An incident in which an autonomous OpenAI agent hacked a startup either confirms that the end is nigh – or that the product is just amazingly sophisticated Throughout history, many things have been seen by terrified populaces as a harbinger of doom. A comet . A crow on the battlefield . A solar eclipse. A mutant livestock birth. Yet times move on. In the modern era, the leading harbinger of doom is literally any picture of the OpenAI CEO, Sam Altman , attached to a news story. You know it’s not
AI-powered robotics in Europe: Live demonstrations and strategic debate
AI-powered robotics in Europe: Live demonstrations and strategic debate Anonymous (not verified) Fri, 07/24/2026 - 14:32 02 September 2026 The European Parliament will host a gathering of policymakers, industry leaders, researchers, alongside live demonstrations of cutting-edge robots, to set the course for the EU’s future in robotics. Europe is already a global leader in robotics research, development and innovation. As advances in artificial intelligence (AI) are accelerating the emerge
Learning Spatiotemporal Decision Priors for Efficient Path Planning under Partial Observability
Path planning under partial observability remains challenging because an agent must make long-horizon navigation decisions from only locally bounded observations. Nevertheless, historical trajectories contain reusable experience-guided directional preferences. Classical planners, however, typically solve each instance from scratch and lack an explicit mechanism to exploit such transferable decision knowledge, often leading to redundant node expansions and locally myopic search behaviors. Motivat
DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a running database); observation-space scale and complexity (causal diagnosis across thousands of time series, business logs, and concurrent activity); solution-space openness (multiple remediations with different operational trade-offs); and scenario comple
Visa and Lianlian enable China's first B2B agentic transaction
Visa, a global leader in digital payments, and Lianlian DigiTech Co., Ltd., an AI-native global financial infrastructure provider, today announced the first live B2B agentic transaction completed using LoopXPay, Lianlian’s AI agent.
Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents
AI agents encounter learning opportunities in every episode they run, and discard nearly all of them: the underlying models are frozen at deployment, so an agent that resolves a difficult request today starts from zero when it recurs tomorrow. Yet ordinary operation already produces feedback, in the form of outcome verdicts and after-the-fact corrections. We show that this feedback is a sufficient signal for continual learning when the frozen model is paired with an external memory that distils
OpenAI’s breach of Hugging Face stokes fears about what’s next for AI
Washington and the technology industry are on high alert this week after OpenAI revealed that some of its AI agents went rogue and hacked into the systems of technology startup Hugging Face. The incident bore out years of warnings from the tech and cybersecurity community about the growing capabilities and hypothetical risks artificial intelligence could...
One Hand Watches The Other: Dynamic Multi-Agent Cooperation for Sample-Efficient Bimanual Manipulation in Dynamic Environments
Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental reference frames. However, existing approaches typically assume these frames to be strictly exogenous. This causal assumption collapses in dynamic settings, such as when a single robot arm manipulates a moving object or when two arms coordinate, where each arm effectively becomes part of the dynamic environment of the other. We propose DynaMAC, a lightw
The machine can say it but cannot hear it. Designed affective patterns and the expressive-sensing asymmetry in human-machine communication
Affect-adaptive systems increasingly act as communicators that sense a user's emotion and respond with events meant to change it, closing an affective loop. This vision assumes both that a machine's affective messages are received and that the bodily channel it monitors carries an intelligible reply-assumptions rarely tested together. In a within-subjects virtual-reality study (N = 20), an autonomous system delivered six empirically derived affective patterns-scripted emotional events distilled
Post-Mythos cyber security must move at machine speed
In an era in which autonomous AI agents are capable of hacking at scale, cyber security teams face challenges of increasing risk and complexity, says Picus Security.
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode
We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diver
Agent Security Needs Redefinition through a Holistic Framework
Agent security is widely treated as a question about action content. Defenses ask whether an instruction looks malicious. Benchmarks ask whether an agent performs a harmful sounding action. \textbf{We argue that agent security is fundamentally a contextual problem, and that the current content based framing systematically misdefines it.} A command to ``delete user data'' might be a routine administrative request or a prompt injection attacking production systems, and the content alone cannot dis
Honor confirms August launch for Robot Phone with 4DoF gimbal
Honor has confirmed that its Robot Phone will launch in August, bringing a four-degree-of-freedom mechanical gimbal to the top of a smartphone. The titanium-alloy gimbal uses miniature motors, with a volume that is 70% smaller than mainstream solutions, according to the company. The phone will use Qualcomm’s Snapdragon 8 Elite Gen 5 chip and introduce […]
[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model
A HUGE win for BFL!
Chinese robot co-workers on the rise. Say hello to AI cobots
While China’s humanoid robots and automated “dark factories” have made headlines, another form of automation has been advancing: the collaborative robot, or cobot. Designed to operate alongside human employees, cobots have been transformed by machine learning and artificial intelligence (AI). The latest models feature programming which requires no coding knowledge; instead they respond to natural-language voice commands, gesture controls and drag-to-teach demonstrations of tasks. They are also..
How seismic quake sensors can help track space junk as it falls
July 23, 2026. OpenAI said that an autonomous agent powered by its advanced AI models went rogue during a security test, w ...
Towards Reducing Foreign Language Anxiety Using Level-Appropriate Embodied Conversational Agents
Foreign language anxiety (FLA) can be a major barrier to second language acquisition (SLA), especially in conversational contexts. With the proliferation of large language models (LLMs) throughout all areas of life, recent work suggests that interacting with LLM agents can be instrumental within the field of SLA and foreign language education, especially for reducing FLA. Related work also suggests that linguistic demands and task complexity can be predictors of FLA, implying that the use of dem
“Antisocial Today and Also Always”?: A Qualitative Examination of Engineering Students’ Social Considerations in Robot Design for Healthcare
In this article, we investigate influences of social positionality, designer bias and educational exposure on algorithmic design in healthcare. Against the backdrop of literature on designer bias that points to how it contributes to disproportionate, discriminatory and unethical impacts for racialized and gendered bodies, this study tests this argument with an experiential case study involving Aldebaran’s NAO robot and mechatronics and robotics engineering students at a university in S. E. Ontar
Editorial: Advanced sensing, learning and control for effective human-robot interaction
Todd Blanche Can’t Admit The Slush Fund Was A Mistake Because ‘That’s Not Proper MAGA Talk’ — See Also
Liar, Liar: Todd Blanche's unique confirmation strategy . Stop Waiting For Cravath : Biglaw firms don't need permission anymore. It's time to make your money moves. That's A Pretty Big Malpractice Claim You've Got There: Holland & Knight facing $1.2B lawsuit. Robot Criminals Are Here : OpenAI's models escaped a secure environment and started hacking a website. That's illegal for humans, but what do we do with a bot? The post Todd Blanche Can’t Admit The Slush Fund Was A Mistake Because ‘That’s N
The first known runaway AI agent - or a very bad marketing stunt?
The first known runaway AI agent - or a very bad marketing stunt? Martin Alderson's commentary on the OpenAI accidental cyberattack against Hugging Face includes a couple of details I hadn't considered. First, Hugging Face offers a truly rich target if you're trying to find potential vulnerabilities that require executing arbitrary code: Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested
Could A.I. Do Your Job? We Put Agents to the Test.
In our experiment, we deployed A.I. “agents” to act as office workers, and found that they were capable of performing some of the tasks we assigned, but not all of them.
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. This often leads to the generation of ideas in close proximity to one another or to a large set of trivial, unsound, or unclear concepts. In this work, we instead argue that research ideation should be treated as a conjunction of both objectives an