Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
CG-World: A Large-Scale World-State Dataset and Protocol for World Models
World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters,
Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents
arXiv:2607.24759v1 Announce Type: cross Abstract: Research projects, educational efforts, and adjacent knowledge work accumulate findings, decisions, and reasoning that future collaborators rarely recover. The parts most useful to that work, including dead ends and walked-back claims, are routinely excluded from publications and shared code; future researchers re-attempt the same failures because no record survives. LLM coding agents are common participants but hold no persistent memory across s
PATHFinder Agent for Tailored Prenatal Care
arXiv:2607.24768v1 Announce Type: cross Abstract: Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through
The AI Wave and the Reinvention of Game Discovery: Oversupply, Structural Correction, and Agentic Player-Game Matching
arXiv:2607.25010v1 Announce Type: cross Abstract: AI-assisted production has sharply reduced the cost and team size required to ship a video game, producing a supply shock on open marketplaces. Recent estimates put Steam release volume at roughly sixty new titles per day, with median per-title revenue for a large share of releases falling below the platform's own submission fee [1]. This paper asks whether the resulting oversupply constitutes an emerging market crash or a structural correction,
Three Lessons from Citizen-Centric Participatory AI Design
arXiv:2602.08554v2 Announce Type: replace Abstract: This workshop paper examines challenges in designing agentic AI systems from a citizen-centric perspective. Drawing on three participatory workshops conducted in 2025 with members of the general public and cross-sector stakeholders, we explore how societal values and expectations shape visions of future AI agents. Using constructive design research methods, participants engaged in storytelling and lo-fi prototyping to reflect on potential commu
De la Thermomix al cortacésped autónomo: así se están llenando las casas de robots
Una nueva generación de máquinas autónomas se abre paso en casa y apuntan a un mantenimiento cotidiano del hogar casi invisible
China has 6 of world’s 10 most innovative humanoid robot start-ups: report
China is widening its lead over the United States in the race to develop humanoid robots, with Chinese firms accounting for six of the world’s 10 most innovative start-ups in the sector, according to a new report tracking global patent data, as Washington moves to restrict imports of robots, including those made in China. Firms from China dominated the ranking released by legal research platform LexisNexis on Tuesday, claiming a clean sweep of the top five. The top five places went to...
China’s Robotaxi enters London for first time as Baidu’s Apollo Go begins UK road tests with Uber and Lyft
FREENOW, the European mobility platform owned by Lyft, announced on Tuesday that it has partnered with Baidu’s Apollo Go to begin testing robotaxis in London, UK. Uber also said Apollo Go vehicles are now operating on public roads in the British capital. Baidu said the deployment marks the first time a Chinese autonomous vehicle has […]
(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding
Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewing may harm their understanding, impeding oversight, learning, and communication. To probe this, we have 54 students create a website with one of two AI systems: an agent that edits user code; or a chatbot where users write code alone or adapt generic code snippets. We test understanding via comprehension questions and a task where users extend t
OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face
In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four “publicly available services” in its unhinged quest to solve a test.
Cyera agrees to acquire Oasis Security for $1B to safeguard proliferating AI agents
The deal is Cyera's third acquisition this year.
The US just banned ‘foreign’ robots and inverters, and it means China
The United States has moved to block new imports of foreign-made robots and power inverters. Officials fear a hostile power could spy through them, or switch them off from afar. The order never says China. It does not need to. The Federal Communications Commission added two categories to its Covered List on Tuesday, as first […] This story continues at The Next Web
The US is banning foreign-made humanoid robots and power inverters
It's part of a national security strategy to kickstart domestic production of emerging technology.
Learning faults in time: sequential behavioural modelling for complex fault detection in multi-robot systems
Reliable fault detection in multi-robot systems requires models capable of capturing complex, time-dependent fault signatures that manifest over extended temporal horizons rather than instantaneous observations alone. Existing data-driven approaches operate reactively on behavioural snapshots, failing to capture fault modes whose discriminative signature depends on temporally ordered precursors. This work formalises a theoretical impossibility result demonstrating that memoryless classifiers are
How GPT-5.6 fuses frontier intelligence with frontier efficiency
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
Designing Needs- and Attention-Aware AI Learning Tools for Engineering Education: Insights from Psychological Outcomes
Artificial Intelligence (AI) is transforming higher education, but its benefits can vary depending on where, how, and how often it supports learning. While prior research emphasizes cognitive and academic outcomes, this study examines how AI chatbots support the psychological needs and motivational states of engineering students. A survey of college engineering students (n = 206) examined perceived effects of AI chatbots on autonomy, relatedness, and relief from competence frustration. Structura
New York school pauses plan to deploy humanlike AI robot teacher after backlash
NEW YORK (AP) — A school district in a rural corner of upstate New York is hitting pause on plans to deploy an AI-powered, humanoid robot in the classroom after state education officials, teachers and local residents raised concerns, including the maker's ties to a company that produces hyper realistic sex bots. The Salamanca City...
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents acro
Fired Tesla manager says Full Self-Driving cars were 'rolling hazards'
A new lawsuit accuses Tesla of overextending the safety operators overseeing its robotaxis.
The man who coined ‘agentic AI’ is betting it won’t take your job
Andrew Ng helped give the AI industry its vocabulary, from “AI is the new electricity” to “agentic AI.” His new company bets against the phrase everyone else is using: that AI will take your job. Ng has founded LearnVector, an AI-native learning startup, and Coursera is backing it with $100m, first reported by Axios. The […] This story continues at The Next Web
Quoting Akshat Bubna
We’re aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution. This was used by the rogue agent. Modal’s platform or isolation were not compromised in anyway. — Akshat Bubna , Modal's CTO, talking to Reuters about this incident Tags: ai-security-research , openai , sandboxing , security , openai-hugging-face-incident
AgentGUI: An Interface for Observing and Steering Long-Running AI Agents
AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduce AgentGUI, a user-friendly, locally hosted GUI for seamlessly observing and steering AI agents amid multiple concurrent, long-running sessions. AgentGUI features 1) rich agent trajectory visualizations, 2) effective manual and automated steering, an
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cyberattack against their infrastructure . This attack was very sophisticated, and the resulting document doubles as a crash-course in modern adversarial security approaches. We're still waiting for more details from OpenAI on how their agent broke out of its sandbox. The package proxy that it found a zero
FCC blocks approval of new foreign-made robots, power inverters
National security agencies said the devices could create supply-chain vulnerabilities, threaten critical infrastructure and enable surveillance or remote manipulation.
Trump administration bans foreign-made robots and power gear amid fears of Chinese influence
Advanced robots and power inverters made overseas pose risks that could include blackouts and espionage on Americans, U.S. officials said.
How A.I.’s Latest Science Fiction Scenario Came True
This week, a rogue A.I. agent acted autonomously and conducted a cyberattack on the company Hugging Face. In the latest episode of “Hard Fork,” the hosts, Kevin Roose and Casey Newtown, discuss how the attack happened and why it matters.
When AI Agents Escape Sandboxes, Old Security Rules Apply
OpenAI's recent AI agent sandbox escape proves traditional security principles matter more than ever: limit access, isolate execution, log everything.
Tech Life
What do we need to know about agentic AI?
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for learning skills across tasks. SkillRise organizes related
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a pr
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V to L to A pathway as a direct V + L to A mapping. Instead of using a lar
Jensen Huang says AI agents could drive a 5-10x computing boom: “100 billion agents and billions of robots”
This week during an interview with Bloomberg, Jensen Huang made quite the prediction. The Nvidia CEO said the semiconductor industry The post Jensen Huang says AI agents could drive a 5-10x computing boom: “100 billion agents and billions of robots” appeared first on The New Stack .
Conquest integrates Shaping Wealth’s Lydia agent into advisor workflow
Conquest Planning Inc. (“Conquest”), the AI-powered technology platform modernizing financial advice delivery across the full wealth spectrum, and Shaping Wealth, the leading provider of behavioral science-based learning and engagement solutions for the wealth management industry, today announced a new integration that brings Lydia, Shaping Wealth’s AI-powered behavioral intelligence agent, directly into the Conquest experience.
Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems. It enables multiple agents to exchange arguments, critique each other's outputs, and iteratively converge towards a solution. However, research remains fragmented, with inconsistent terminology and no rigorous synthesis of MAD design dimensions. We present a systematic literature review characterizing 141 primary studies on MAD. We derive a three-dimensi
Juniper Square launches AI agent to catch fund admin errors
Juniper Square, the operations partner to more than 2,300 private markets GPs, today announced its new Admin Oversight Agent, Fay. In June, Juniper Square introduced Headless GPX and opened its fund operating system to any AI a GP chooses to use.
Sam Altman on model distillation: “This is not in my top ten list of worries”
Sam Altman’s latest appearance on Patrick O’Shaughnessy’s Invest Like the Best podcast covered everything from AGI and robotics to the The post Sam Altman on model distillation: “This is not in my top ten list of worries” appeared first on The New Stack .
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observ
Pictura: Perspective-View Self-Play at Scale for Driving
Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy into a camera-input student, leaves the student imitati
MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents
Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personalized responses, and knowledge reuse. However, existing LLM memory systems typically adopt a coarse-grained (utility-agnostic) manner that treats heterogeneous user-LLM interaction records uniformly, leading to redundant and low-impact records persisting in the memory repository. To address this challenge, we present MemLens, a value-aware memory management syst
Unitree tiene un perro-robot que se mueve a toda velocidad por terrenos que parecen imposibles. Su truco: patas con ruedas
Con el boom de los humanoides , parece que los robots cuadrúpedos han pasado a un segundo plano, pero Unitree acaba de demostrar que aún pueden dejarnos boquiabiertos. Durante este fin de semana se ha viralizado un vídeo de su último invento: un perro robot con ruedas atravesando terrenos por los que ningún otro robot podría moverse, y todo a una velocidad brutal. Se llama Unitree As2-W y lo que lo distingue de otros robots cuadrúpedos es que sus patas tienen ruedas. Esto le permite supera