Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
Least privilege for AI agents: Identity, access, and tool binding
As AI agents become more autonomous, strong identity, access, and auditing controls are critical to keeping them secure.
MeitY proposes mandatory human-in-the-loop interventions in agentic AI payments
CERT-In has proposed mandatory human oversight for high-value agentic AI payments, as NPCI and fintech firms develop protocols that could allow AI agents to make UPI transactions. The post MeitY proposes mandatory human-in-the-loop interventions in agentic AI payments appeared first on MEDIANAMA .
Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction
As conversational AI systems are designed for repeated use, a central question is how a series of interactions becomes a relationship. We present a longitudinal multimodal study of a memory-augmented conversational agent (24 participants x 10 sessions), in which participants rated five relational constructs -- familiarity, self-disclosure, perceived memory, conversational quality, and enjoyment -- after each session. Two complementary dynamics emerge. First, conversational quality strongly shape
SESAME 2026 : 2nd Workshop on Smarter Extraction of ScholArly MEtadata using Knowledge Graphs, Language Models and Agents (SESAME) at JCDL 2026
2nd Workshop on Smarter Extraction of ScholArly MEtadata using Knowledge Graphs, Language Models and Agents (SESAME) at JCDL 2026 [Dallas, Texas, USA] [Oct 13, 2026 - Oct 26, 2026]
Democratizing Agent Deployment Safety: A Structural Monitoring Approach
AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an agent completes its assigned task while covertly weakening safeguards through actions such as broadening permissions, degrading logging, or introducing persistence mechanisms. While frontier laboratories may deploy sophisticated monitoring pipelines, many organizations and individual users adopting coding agents lack the resources and governance maturity requ
Robots, AI and drones: how the Dutch navy is using tech to transform its sea defences
Uncrewed systems are the future for armed forces and the Netherlands is leading the way ‘to keep people out of danger zones’ On each side of the target ship, a black vessel keeps a watchful distance. Defender 1 and Defender 2 are the eyes and ears of the navy – but they have nobody onboard, and their paths are controlled by a computer system. This is the future of the Royal Netherlands Navy, according to Capt Sjoerd Feenstra, head of the expertise centre for unmanned systems. He is leading a fiv
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
arXiv:2607.13041v1 Announce Type: new Abstract: Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them. This study introduces LessonBench-V1, a benchmark dataset comprising 647 human-written lessons paired with LLM-based reverse-engineered lesson plans across 240 STEM topics spanning mathematics, physics, chemistry, and computer science. The lessons are drawn from 97 trusted
The tragedy of the cognitive commons: collective intelligence beyond AI-induced knowledge collapse
arXiv:2607.13272v1 Announce Type: new Abstract: In a recent dynamic model by Acemoglu, Kong and Ozdaglar (2026a) agentic AI can cause a self-reinforcing deterioration of humanity's common knowledge base in what they call knowledge collapse. The model is based on a natural complementarity between the cumulative general knowledge of humans and locally generated context-specific knowledge, and on a learning externality that means that we all contribute to the private signal and the thin public sign
Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
arXiv:2607.13370v1 Announce Type: new Abstract: This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledge Component (KC) models across integrated Chat, Tutor, and Quiz modes. That prior work validated LEA on a single STEM course (CMP511) exclusively through simulation, using synthetic learner agents. This
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
arXiv:2607.13230v1 Announce Type: cross Abstract: Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interact with third-party services. This paper develops an AI-native mathematical framework for underwriting, pricing, and contract design for agentic AI deployments. A deployment is represented by a risk state that captures autonomy level, operational authority, permission exposure, governance maturity,
Design of policy digital twins incorporating multi-level agent based modelling
arXiv:2607.13766v1 Announce Type: cross Abstract: Digital twins are used across many industries to enable better decision making. However, while policy makers at all levels (including city, national and supranational scales) have expressed a desire to integrate digital twins into their workflows, this adoption has been slow to materialise. In this paper, we discuss the key issues associated with policy digital twins, and the ways in which they differ from, and are similar to, their counterparts
Early Adoption of Agentic Coding Tools by GitHub Projects
arXiv:2607.14037v1 Announce Type: cross Abstract: Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level outcomes of agent-generated contributions, less is known about how agentic coding tools are adopted and managed at the project level. In this paper, we analyze 25,264 agentic PRs from 2,361 popular GitHub repositor
The Agentic Web Requires New Normative Infrastructure
arXiv:2606.10711v2 Announce Type: replace Abstract: The agentic web, in which users interact with the internet largely through agents acting on their behalf, is now technically feasible. However, many of the consumer and social benefits that could be realized by online AI agents acting scrupulously in their principals' interest are currently obstructed by outdated laws, terms of service, and other less formal practices which allow online platforms to block and degrade agent access, often in secr
Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation
arXiv:2606.02528v2 Announce Type: replace-cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. We ask three questions: do LLMs systematically prefer certain financial instruments; can an internal representation with causal leverage over those preferences be identified; and does that representation affect downstream financial decisions? We develop a three-level audit protocol and apply
Three insights you may have missed from theCUBE’s coverage of RAISE Summit
Agentic inference is reshaping the center of gravity in AI infrastructure. What began as a race to scale training has shifted into a phase defined by expanding context windows, memory‑augmented reasoning and the need to keep graphics processing units continuously fed with data. As enterprises push deeper into agentic systems, storage has moved into the […] The post Three insights you may have missed from theCUBE’s coverage of RAISE Summit appeared first on SiliconANGLE .
Xiaomi introduces Xiaomi-Robotics-U0 for embodied AI and robot generation
Xiaomi on Wednesday unveiled Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive foundation model for embodied AI. The company said the model unifies four core capabilities within a single framework: embodied scene generation, embodied transfer, robot interaction video generation, and general-purpose image generation and editing. According to Xiaomi, the model can generate robot-ready environments from text prompts, adapt […]
Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers
AI assistants increasingly retrieve web content at inference time to provide fresh and grounded answers, yet it remains unclear whether these search-augmented capabilities respect website-owner restrictions expressed through robots.txt. We present a controlled empirical study of ten widely used AI assistants with advertised web-search capabilities. For each assistant, we first identify a configuration that actually produces observable web-browsing behavior and record the user-agent exposed durin
Mermaid to Unicode box art (grok-mermaid)
Tool: Mermaid to Unicode box art (grok-mermaid) While exploring the codebase for the newly open-sourced Grok CLI coding agent I came across xai-grok-markdown/src/mermaid.rs , a "self-contained terminal renderer for Mermaid diagrams" written in Rust. I figured it would be fun to try that out in a browser via WebAssembly. Here's the prompt I ran in Claude Code for web (Fable 5), and this is what the resulting tool looks like: Tags: tools , rust , webassembly , mermaid , grok , xai
Mastercard picks UK as test environment for agentic AI sandbox
Mastercard is opening an agentic AI sandbox in the UK, to help banks and retaillers explore, test and validate use cases before introducing them into production
How Cars24 scales conversations and builds faster with OpenAI
Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams across the company.
Correction: Synthetic chamber: agentic mediation in representative democracy
Exploring generative AI through core frameworks, emerging innovations, and applications
Generative Artificial Intelligence has undergone rapid maturation between 2023 and 2025, driven by three converging paradigm shifts: the emergence of multimodal foundation models unifying text, image, audio, and video synthesis; the rise of agentic autonomy transforming generative systems into goal-driven, autonomous entities; and the formalization of responsible AI governance through legally enforceable regulations. While this technological landscape has generated substantial economic impact, c
Hybrid task and motion planning with reactive collision handling for multi-robot disassembly of complex products: application to EV batteries
This paper addresses the problem of multi-robot coordination for complex manipulation task sequences. We present a vision-driven task-and-motion planning (TAMP) framework for a real dual-agent platform that integrates task decomposition and allocation with a learning-based planner. A GMM-informed RRT motion planner is coupled with a hybrid safety layer that combines predictive collision checking in a MoveIt/FCL digital twin with reactive avoidance and replanning. This integration is challenging
Structural predictors and latent maturity regimes of robotic readiness in global health systems: evidence from machine learning-based latent clustering and class prediction
BackgroundThe systematic integration of robotics into health service delivery systems requires periodic assessment of robotic readiness in terms of digital-health maturity regimes across countries. The current study aims to cluster 169 countries into maturity regimes and classify and predict cluster membership accuracy based on digital-health maturity dimensions determining the system’s perception and interoperability, coordination, and workforce–regulatory reliability readiness. These country-l
NVIDIA Introduces New Jetson Thor Computers to Advance Mainstream Robotics and Edge AI
General-purpose robots and autonomous machines are moving from research labs to real-world mass-market deployment, creating demand for compact, power-efficient AI supercomputers capable of running foundation models at the edge. To meet that need, NVIDIA today introduced the T3000 and T2000, new modules based on the NVIDIA Thor architecture that enable mass-market robotics and edge AI […]
Cadence extends its AI agents beyond chips with AuraStack for circuit boards and packaging
Integrated circuit and electronic hardware design company Cadence Design Systems Inc. today announced a new artificial intelligence agent that assists with packaging and system design, the step after silicon has been customized and produced, marking how chips and components get used in electronic devices. AuraStack, a super-agent within the company’s Allegro AI studio, provides agentic […] The post Cadence extends its AI agents beyond chips with AuraStack for circuit boards and packaging appeare
Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents
Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agents” are still chatbot wrappers, the control plane enterprises expect is deliberately hybrid to avoid lock-in, and real-time fiscal control over token burn remains the exception. This wave of VentureBeat
An offline approach to fNIRS-guided reinforcement learning for robot behavior
Human-in-the-loop Reinforcement Learning has become a popular approach to training, finetuning, and aligning robot behavior with user preferences. Our paper explores the feasibility of using brain signals via functional near-infrared spectroscopy (fNIRS) to modulate robot learning in simulation. We compare agents trained on passive (observational) versus active (demonstrative) interaction tasks, and test multiple methods for enhancing the RL algorithm with the neural signal, focusing on paramete
AWS adds AI-assisted product listing service to its Marketplace portfolio
Product listings in AWS Marketplace gained new AI-based features last month in anticipation of continued growth in the use of enterprise agents. Amazon Web Services Inc. unveiled AI-assisted product listings in Product Assistant chat, a feature that helps independent software vendors and consulting partners develop comprehensive product listings for AWS Marketplace using existing digital assets. […] The post AWS adds AI-assisted product listing service to its Marketplace portfolio appeared first
When Is Delegated Play Truthful? Within-Range Regret and the Trilemma of Aligned Delegation
Advertisers delegate bidding to autobidders; users delegate tasks to language-model agents. A person describes what they want to an automated proxy that acts in a mechanism on their behalf. This is the revelation principle in production, and it forces a question classical theory assumes away: when is it optimal to describe yourself honestly to your own proxy? We show the answer turns on one quantity, the proxy's within-range regret. The most a principal can gain by misreporting equals the regret
Boston Dynamics is testing Spot as a package delivery robot that walks from van to doorstep
Boston Dynamics is testing Spot, its four-legged robot, as a delivery worker that rides along in trucks and hops out to carry packages to customers’ doorsteps. The company posted a video on Tuesday showing a human driver loading packages onto a conveyor belt mounted on Spot’s back, followed by the robot walking up to houses […] This story continues at The Next Web
Walden Robotics launches with $300M, and its factory humanoids have no legs
Walden Robotics has launched from stealth with $300m and a $1.1bn valuation. The Toyota spin-out builds humanoids for factories, and it left off the legs on purpose. “If you listen to the people on the factory floor, they aren’t ready, and they don’t want them yet.” That is Russ Tedrake, telling Bloomberg why his new […] This story continues at The Next Web
Construction robot startup Monumental reels in $32M
Monumental BV, a Dutch startup that operates a fleet of bricklaying robots, has raised $35 million in funding. Khosla Ventures led the Series B investment. Monumental stated in its funding announcement today that existing backers Plural and Hummingbird chipped in as well. Monumental uses robots to construct the walls of homes, schools and other buildings. […] The post Construction robot startup Monumental reels in $32M appeared first on SiliconANGLE .
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can become trapped in repetitive loops, wasting search budgets and ultimately compromising the quality and completeness of the final output. We introduce SearchOS, a system-level multi-age
BadWAM: When World-Action Models Dream Right but Act Wrong
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluatin
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between episode-level outcomes and token-level policy learning. We propose SEED (SElf-Evolving On-Policy Distillation), a self-evolving
RoboTTT: Context Scaling for Robot Policies
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturb
SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment
CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to CAD models, yet typically their correspondences are appearance-driven and degrade under occlusion or sim-to-real domain shift. To address these limitations, we introduce SUFLECA (Scaling Up Feature LEarning for CAD Alignment), a we
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment. The gap is especially important for AI agents, whose observations, tool outputs, documents, and prior decisions accumulate over long trajectories. LongStraw is an architecture-aware execution stack for million-token RL post-training under a fixed GPU bu
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by