Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
In a world of AI agents, where do we fit in?
For more than a decade, leaders have used the phrase “Future of Work” to describe how technology is transforming business. The post In a world of AI agents, where do we fit in? appeared first on The New Stack .
A Fireside Chat with Cat and Thariq from the Claude Code team
Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves. The full video of the session is now available on YouTube . Below is an edited copy of the transcript, with extra links and my own bolded highlights. A few top-level notes if you don't want to watch the video or
Deep learning-based prediction of time-resolved adhesive forces in viscoelastic Hertzian contacts
Fast prediction of the response of adhesive soft viscoelastic contacts represents a current challenge in soft robotics and for gripping and manipulation tasks. Determining the complete time-resolved force trajectory requires full numerical simulations, whose computational cost is strongly parameter-dependent, making them impractical for real-time application or design-optimization loops. In this work, we overcome this limitation by training a scalar-conditioned, stateful, sequence-to-sequence de
FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives. While current video generation models excel at short, single-scene clips within narrow temporal and spatial contexts, novel-to-film generation operates in a more complex regime, demanding long-duration content across diverse scenes with dynamically evolving entity states. To address this, we formalize novel-to
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a
Native Acquires Frontline Research Group to Build the Agentic AI Data Layer for Africa’s $1.7 Trillion Traditional Trade Market
Native, the agentic intelligence company building the operating system for offline trade, today announced it has acquired Frontline Research Group, a leading market intelligence business serving consumer goods companies across more than 14 African markets.
In Straße mit Gefahrenstelle eingefahren: Zoox ruft Robotaxis zurück
Weil ein Robotaxi Rauch ignoriert und sich einer Brandstelle genähert hat, ruft Zoox seine Fahrzeuge zurück. Laut einer Verkehrsbehörde war es kein Einzelfall.
Gritt exits stealth with $32 million for robots to build solar plants — then, everything else
Gritt is coming out of stealth with $34 million and plans to automate the hardest tasks on construction sites.
The coming burnout from managing AI agents
Technology leaders are very comfortable talking about productivity . We can discuss automation rates, developer velocity, cost curves, and defect rates with great confidence. We are less comfortable talking about mental strain, especially when the thing creating the tension is the very thing we are excited to adopt. Artificial intelligence agents will make that discomfort harder to avoid. Agents can write code, triage issues, research options, test hypotheses, and perform many of the tasks that
Humanoid robots set for explosive growth
The robots are expanding into healthcare, logistics, retail, hospitality and education as AI advances fuel global demand.
The next enterprise AI frontier is the optimizable company
For the past two years, companies have been trying to add AI to their processes. That sounded reasonable at first. Existing workflows, existing systems, existing data, existing dashboards — now with intelligence attached. Add a copilot here, an agent there, a few automated steps, a model connected to some tools, and wait for productivity to rise. But this approach is starting to reveal its limit: The problem isn’t that AI can’t act. It increasingly can. The problem is that most companies don’t h
Xiaomi-Robotics-1 shows that more data beats bigger models when training robots to move
Xiaomi trained Xiaomi-Robotics-1 on more than 100,000 hours of motion data collected by people using camera-equipped handheld grippers rather than robots. Adding data improved performance far more than increasing model size. The gains haven't plateaued, though absolute success rates remain low. The article Xiaomi-Robotics-1 shows that more data beats bigger models when training robots to move appeared first on The Decoder .
Attaquée par un agent IA autonome, Hugging Face a analysé les traces avec un LLM local
Hugging Face a publié le 16 juillet 2026 une divulgation d’incident au sujet d’une intrusion dans une partie de son infrastructure de production. Selon l’entreprise, cette intrusion présentait une caractéristique inédite : elle a été pilotée de bout en bout par un système d’agent IA autonome. Au constat de cette intrusion, Hugging Face en a […]
Data Leakage Prevention in Agentic Applications via Preemptive Hardening
Agentic systems integrate LLM driven planning with interfaces to external tools, making data leakage and tool misuse feasible via instruction/data boundary failures and prompt injection attacks. Enforcing required controls consistently is particularly challenging in workflows spanning many codebases and heterogeneous agents. To address this challenge in multi agentic systems, we present a pre-deployment pipeline for scanning, hardening, and validation of agentic applications. The pipeline analyz
BrainCo demos thought-controlled robots at WAIC 2026
BrainCo, a Hangzhou brain-computer interface developer, demonstrated a brain-controlled robot platform at WAIC 2026. The system uses an EEG headset to translate a user’s neural signals into commands for machines. The company says the platform can connect to humanoid robots, robotic arms and robot dogs. It is designed to translate imagined actions into physical operations […]
AgentTrails: Towards Trust and Reuse for Agentic Tasks
LLM-powered agents increasingly tackle complex tasks by invoking tools, querying databases, executing code, and manipulating intermediate artifacts. These agents follow trajectories that are typically stored as chronological logs, obscuring the underlying dataflow -- the dependencies between their actions and the artifacts they create and manipulate. This limits developers' ability to understand the agents' trails, compare executions, debug failures, and re-use the computations. We present Agent
XPENG releases TuringViT for smart driving and humanoid robots
XPENG has published TuringViT, a vision encoder for vision-language and vision-language-action models used in smart driving, smart cockpit systems and its IRON humanoid robot program. XPENG offers TuringViT-18L and TuringViT-24L. At 1536×1536 resolution, the company says TuringViT-18L reached 3.04 times the throughput of Seed1.5-ViT and 2.16 times that of SigLIP2-ViT-L. XPENG says the model was […]
KI-Training: Tesla l�sst Gr�nheide-Belegschaft f�r Roboter Optimus filmen
Ausgew�hlte Tesla-Teams sollen Kameras am Arbeitsplatz tragen. Die Aufnahmen sollen dem Roboter Optimus beibringen, wie Menschen arbeiten. ( Gigafactory Berlin , Roboter )
Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
AI-native biotechnology companies are often designed by copying human biotech org charts into agent roles. We argue for a different abstraction: a Company World Model, defined as a persistent asset-to-value state representation with transition models, explicit value functions, planning, and updating across scientific, regulatory, BD, commercial, financial, and execution constraints. We introduce a dry-lab benchmark for testing whether AI-agent organizations should mimic departments or operate ar
In Graphic Detail: AI visibility is no longer about referral traffic
New data shows why publishers are developing strategies around agentic traffic and brand visibility in AI answer engines.
A Diagnostic Framework for AI Agent Behavior
arXiv:2607.17149v1 Announce Type: cross Abstract: AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer att
The Optimization Trilemma: Efficiency, Comfort and Fairness in Decentralized Multi-agent Coordination
arXiv:2607.17311v1 Announce Type: cross Abstract: The problem of fair multi-agent coordination in decentralized settings is one of the most pressing challenges for building efficient collaborative systems. Resource allocation is based on optimized collective arrangements accounting for agents' needs. Such coordination should not only be computationally efficient but also account for fairness, i.e., equitable redistribution of costs incurred by all agents. Recent literature has proposed several a
The atomic structure of work: a micro-action instrument reveals two-pole AI occupational exposure and its decade-scale polar inversion
arXiv:2606.07939v2 Announce Type: replace Abstract: Research on artificial intelligence and work assigns each occupation a single exposure score. We build an instrument to see what those scores average over: a decomposition of 1,961 O*NET work activities into 15,817 atomic micro-actions by a consensus multi-agent LLM pipeline, clustered from text alone into seven semantic classes. Projecting exposure indicators onto these classes reveals two extreme poles, tool-mediated physical execution and pl
Artificially intelligent agents in the social and behavioral sciences: A history and outlook
arXiv:2510.05743v3 Announce Type: replace-cross Abstract: We review the historical development and current trends of artificially intelligent agents (agentic AI) in the social and behavioral sciences: from the first programmable computers, and social simulations soon thereafter, to today's experiments with large language models. This overview emphasizes the role of AI in the scientific process and the changes brought about, both through technological advancements and the broader evolution of sci
AI Contagion in Social Networks
arXiv:2606.15206v2 Announce Type: replace-cross Abstract: We study how artificial intelligence (AI) interacts with social communication networks to shape the stability of collective knowledge. Agents exchange information through a network while AI systems generate content and retrain on the aggregate informational environment they influence. This interaction creates a recursive feedback loop in which informational distortions diffuse through society and subsequently feed back into future AI outp
Grabette: an open system to record robot-manipulation data
Environment-free Synthetic Data Generation for API-Calling Agents
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method genera
At World AI Forum, Four Signs China’s AI Industry Is Growing Up
This year, China’s biggest AI conference showed an industry racing to cut computing costs, put AI agents to work, manage new risks, and find business models that can last.
‘They’re going to kill me’: A Waymo rider was trapped inside as vandals smashed the robotaxi
As the Waymo pulled up to a stoplight in San Francisco's Marina District, Sherman Watson raked his eyes over the empty road and saw a silhouette approaching in the darkness. He was riding shotgun, peering out at the vast grid of Pierce and ... (https://incidentdatabase.ai/cite/1599#7531)
Zoox Recall Affects 105 Robotaxis After Fire Scene Incident
Zoox has issued a recall of the software driving its fleet of 105 robotaxis after one of its autonomous vehicles drove into a smoke-obscured emergency fire scene in Las Vegas, the latest incident to draw federal scrutiny over how driverless ... (https://incidentdatabase.ai/cite/1602#7535)
Chinese robot makers’ lament: if we only had a better ‘brain’, and more data
Chinese robotics companies lack both sufficient data and a good “brain” to improve the interaction of their products with the physical world, according to industry insiders at the World Artificial Intelligence Conference (WAIC), which concluded on Monday in Shanghai. The most critical challenge for the embodied AI industry was to “link hardware, data, models and real-world scenarios into a closed-loop iterative system”, said Wang Xiaogang, co-founder of SenseTime and chairman of its robotics...
State of Data & AI 2026
Scaling AI, Agentic AI and Data Sovereignty & Compliance.
ChatMuse: Supporting In-Person Small-Group Conversation Experience with a Proactive Assistive AI Agent in Mixed Reality
In-person small-group conversations occur across nearly every aspect of daily life and play a crucial role in social interaction. However, achieving effective in-person group conversations can be challenging and cognitively demanding. While recent Mixed Reality (MR) headsets show promise as a conversational support system by presenting relevant information through overlays, it remains unclear how such supporting information should be designed and generated for in-person group conversations. We p
MAGE: Human-Like Macro Placement via Agentic Multimodal Reasoning
Macro placement still requires substantial manual refinement in industrial physical design flows. We present MAGE (Macro Placement Agentic Engine), a multimodal multi-agent framework for macro placement refinement. MAGE decomposes the macro placement task into a six-phase workflow that combines structured floorplanning rules, visual checks, and iterative refinement. Expert floorplanning knowledge is encoded through natural-language directives and validation criteria, rather than learned from lab
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable a
Masked Visual Actions for Unified World Modeling
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrar
NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipeline, and the resulting task distributions often reflect substrate biases rather than real-world demand. We introduce NexForge, a requirement-driven framework that takes high-level capability requirements as input and synth
AutoIndex: Learning Representation Programs for Retrieval
We present AutoIndex, a framework for learning representation programs: executable transformations that map raw documents into the representations exposed to a retrieval system. Rather than tuning retrievers, rerankers, or a small set of preprocessing hyperparameters, AutoIndex searches over programs that slice, enrich, normalize, reweight, or reorganize documents before indexing. At each iteration, AutoIndex performs validation-guided program search, in which agents diagnose failures of the cur
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global traject
State of Data & AI 2026: Agentic AI
Journey Beyond finds safer path to customer AI with agentic agents.