13:52 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

Exploring deep reinforcement learning acceleration by superscaling data augmentation via branched fractal symmetries

Learning deep reinforcement learning (DRL) policies directly in physical robots remains bottlenecked by slow wall-clock training times. We present preliminary research on Branched Euclidean Group Fractal Symmetries, a trajectory-level augmentation framework that super-scales group transformations to accelerate policy learning for manipulation. We model a Markov decision process (MDP) as a tree of state–action pairs; at each depth, affine transformations generate geometric structures, within whic
Frontiers in Robotics and AI 31d ago Research RegulationAgents & autonomy

Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows

LLMs increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidence that triggered a repair. We introduce Agentic Transaction Processing (ATP), a transaction model that treats generated actions as untrusted proposals until they pass deterministic admission under a declared, executable constraint set C. The governing principle is two-sided: a proposal is not truth, and no proposal for
arXiv 31d ago Research Agents & autonomy

Fake Bug Report Hijacks AI Coding Agents at Scale

"Agentjacking" is the latest demonstration of how easily attackers can exploit an AI agent's inability to differentiate between content and instructions.
Dark Reading (AI security) 31d ago News Agents & autonomy

Would You Marry Superintelligence?

Emotional bonds between humans and AI companions are growing, and the question of whether a person may marry an AI system will soon move from speculative fiction into law. This chapter examines whether the autonomy-centered logic that has expanded marital choice among human beings can justify extending marital status to superintelligent companions. Following a scenario-envisioning exercise informed by anticipatory ethics, I argue that granting such status leads to socially unjust outcomes, even
arXiv 31d ago Research RegulationAgents & autonomy

ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

Hugging Face Blog 31d ago Field notes Agents & autonomy

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-only rewards provide too sparse guidance, failing to inform the model about the goodness of intermediate actions. Dense supervision methods aim to solve this problem by scoring intermediate steps, from intrinsic confidence to self-distillation and embedding similarities. However, it is common practice to evaluate them by measuring the downstream perfo
arXiv 31d ago Research Agents & autonomy

The Shifting Fortunes of the Kurds

The Kurds’ fortunes have ebbed and flowed in recent years, but the fall of the Assad regime in Syria in December 2024, the 2025 decision by the Kurdistan Workers’ Party (PKK) to dissolve and engage in talks with the Turkish government, and the 2026 U.S.-Israeli war with Iran had enormous ripple effects on the lives of Kurds in the Middle East and Kurdish hopes for autonomy. We asked four experts to assess how recent regional events are presenting risks and opportunities for the Kurds in Turkey,
War on the Rocks 31d ago News Agents & autonomy

NVIDIA BioNeMo Agent Toolkit Brings Accelerated AI to Life Sciences Researchers in Claude Science

Life sciences has entered an era of computational scale, and for more than a decade, NVIDIA has built the full GPU-accelerated computing stack — spanning hardware, frameworks, libraries, models, microservices and domain-specific tools — to help researchers run more sophisticated workflows and iterate faster. This week, Anthropic announced Claude Science, an AI workbench for science […]
NVIDIA Blog (AI) 31d ago Field notes Agents & autonomy

MVP-Nav: Multi-layer Value Map Planner Navigator

Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment. Existing approaches either rely on high-level semantic reasoning without geometric grounding or learn end-to-end policies that lack explicit physical constraints, often resulting in semantically plausible but physically unsafe behaviors. In this paper, we propose
arXiv 31d ago Research Safety & alignmentAgents & autonomy

Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling

A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observable closed-loop environments. In such logs, the ego vehicle has rich local observations, while surrounding agents are only partially observed due to perception limits and occlusions. As a result, simulators may learn incomplete context--action mappings that remain hidden in log-based training but emerge during closed-loop rollouts, leading to unreal
arXiv 31d ago Research Agents & autonomyEnvironment

Q&A: What is agentic AI today, and what do we want it to be?

Computer scientist Phillip Isola cuts through the hype to explain how AI agents work and what the future might hold for this rapidly advancing technology.
MIT News 31d ago News Agents & autonomy

How Jaiveer Singh Is Helping Robots — and Developers — Move Faster

When Jaiveer Singh talks about robots, he doesn’t begin with spectacle. He begins with infrastructure: the boards inside machines, the software that lets developers see through a robot’s cameras and the engineering required before a robot can leave a demo floor to do something useful. As a robotics software engineer who leads the team behind […]
NVIDIA Blog (AI) 31d ago Field notes Agents & autonomy

A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical execution. We developed ProtoPilot, a self-evolving multi-agent system, together with an expert-grounded benchmark and evaluation framework for testing this conversion as an experimental automation problem. The framework spans 294 synthetic-biology and molecular
arXiv 31d ago Research Jobs & economyAgents & autonomy

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, coding environments, robotic systems, security-operation workflows, and autonomous agents that can read private data, call tools, write files, execute code, and act across organizational boundaries. This shift changes the security problem: risks do not arise from the model weights alone, but from the full lifecycle and application stack through which data, promp
arXiv 32d ago Research Agents & autonomyEnvironment

Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning

Editor’s note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners, and enterprises can transform their workflows using the latest advances in OpenUSD and NVIDIA Omniverse. Vision AI agents are becoming a practical way to automatically turn video data from the physical world into operational intelligence in factories, […]
NVIDIA Blog (AI) 32d ago Field notes Agents & autonomy

What Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability

Alignment Forum 32d ago Research Agents & autonomy

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment. In practice, however, as market context accumulates over long horizons, these mandates gradually lose their behavioral influence, a phenomenon we formalize as Mandate Salience Decay (MSD). To measure MSD objectively, we introduce FinPersona-Bench, a
arXiv 32d ago Research Agents & autonomy

Army using AI, robot boats for Pacific logistics

“If you can work in the Pacific, you can work anywhere in the world,” said Maj. Gen. Gavin Gardner.
Defense One Technology 32d ago News Agents & autonomy

Stage-Transition Dense Reward Modeling for Reinforcement Learning

Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations. This work proposes Stage-Transition Dense Reward (STDR), a visual reward-learning framework that converts unstructured expert videos into logically grounded dense rewards for training RL agents from scratch. STDR leverages semantic understanding to infer a task's stag
arXiv 32d ago Research Agents & autonomyEnvironment

Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling

Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry. We present Embodied CAD, solver-grounded LLM agents for parametric B-Rep assembly modeling. Instead of generating a complete script in one pass, the agent iteratively selects actions from a strati
arXiv 32d ago Research Agents & autonomy

Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?

As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing values. However, existing research on LLMs with moral dilemmas overlooks a central aspect of human moral cognition: the ability to imagine alternatives that move beyond the given options. We introduce MoralAltDataset, a dataset of 307 moral dilemmas spanning narrative Advisor dilemmas and AI-facing Agent dilemmas, each augmented with compromise and reframed
arXiv 32d ago Research Agents & autonomy

Long-term Traffic Simulation via Structured Autoregressive Modeling

Interactive traffic simulation is a vital world model for autonomous driving. A central challenge in long-horizon simulation is modeling sustained multi-agent interactions, which is further exacerbated by dynamic token cardinality as agents continuously enter and exit the scene. In this work, we propose that the solution lies in the synergy between the architectural inductive biases and statistical priors of large-scale sequence models, e.g., Large Language Models (LLMs). Our probing experiments
arXiv 32d ago Research Agents & autonomy

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps between high-level instructions and low-level motor commands, and inference inefficiency due to autoregressive scalar decoding. In this work, we propose MIRTH, a unified framework designed to address these challenges. MIRTH
arXiv 32d ago Research Agents & autonomy

Artificial Resonance: AI companions as agents of social acceleration

AI companions are becoming increasingly popular, with millions of users worldwide, especially young adults. Some see the potential to fight the so-called loneliness epidemic; others see the destructive effects of addiction and harmful guidance leading users in extreme cases even to suicide. Recent research has examined AI companions through the lens of AI ethics, addressing questions of emotional dependency, controllability and emotional harm. While these contributions are valuable in assessing
AI & Society 32d ago Research Agents & autonomy

A new paradigm for marine ecological monitoring through swarm intelligence, digital twins, and Human–Swarm interaction

Marine and coastal ecosystems are among the least observable yet most rapidly changing environments, where climate impacts, pollution, and biodiversity loss demand monitoring and intervention at scales that manual sampling and single-robot deployments cannot sustain. This paper argues for a conceptual shift in ecological monitoring and restoration toward networked robotic ecosystems, adopting cooperative swarms of autonomous aquatic robots coupled to in-situ digital twins and human-in-the-loop s
Frontiers in Robotics and AI 32d ago Research Agents & autonomyEnvironment

Low-cost social robot designs for education: a review

Social robots have shown promising potential in educational contexts worldwide, with studies reporting significant cognitive and affective gains when such robots are deployed. However, among other factors, the high cost of commercial robots limits this line of research to a small number of laboratories and hinders large-scale adoption in real-world educational settings, with most studies remaining short-term pilot interventions. Although several reviews exist in this domain, they primarily focus
Frontiers in Robotics and AI 32d ago Research Children & educationAgents & autonomy

The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflows

Agentic artificial intelligence is increasingly deployed not as a single assistant but as a collective of planners, solvers, reviewers, memory managers, tool users, and orchestrators. These systems are entering organisational workflows under familiar labels such as teams, managers, committees, markets, and workflows. This article asks whether such agent collectives exhibit organisational behaviour in a sense that is analytically comparable to, yet distinct from, human organisational behaviour. I
arXiv cs.HC 32d ago Research Agents & autonomy

Behavioral Governance for Autonomous AI Agents: The AgentBound Framework

Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications, and enterprise workflows. Existing agent infrastructure relies on identity federation and delegated authorization to authenticate workloads and control resource access, but it cannot determine whether an authorized action should be executed under the current behavioral and operational context. We present AgentBound, a runtime governance framewo
arXiv 32d ago Research RegulationAgents & autonomy

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric. We introduce a framework that formulates therapeutic response generation as a decision-refinement problem driven by multi-dimensional, human-aligned evaluation. In Stage I, we introduce TheraJudge, an open-source therapeutic evaluator trained via preference-based optimization on human-annotated data to produce
arXiv 32d ago Research HealthcareAgents & autonomy

AI agents are not your “coworkers”

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Imagine coming in to work to learn that a new underling will report to you. The worker is not a person but an AI tool—one that your company nonetheless calls Alex, an…
MIT Technology Review 32d ago News Agents & autonomy

Claude Meets Blackwell Ultra: Anthropic’s Models Now Run on NVIDIA GB300 in Azure

Anthropic’s Claude models in Microsoft Foundry — hosted on Microsoft Azure and running on NVIDIA GB300 Blackwell Ultra GPUs — are now generally available, giving Azure-native enterprises a powerful new way to build autonomous and domain-specific AI agents. As agentic AI continues to drive enterprise innovation and becomes more autonomous, organizations need access to computing […]
NVIDIA Blog (AI) 32d ago Field notes Agents & autonomy

ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis

Multi-agent LLM systems can decompose software-engineering work into planning, generation, validation, and repair, but a narrower systems problem remains: before any governed shared mutation is applied, a system must decide which concurrently formed write intents may proceed in parallel, which require deterministic composition or serialization, and which must take a fail-closed path. We address this problem with the AI-Atomic-Framework (ATM), a specification-grounded governance substrate for sof
arXiv 32d ago Research RegulationAgents & autonomy

Trust Issues Could Make or Break Agentic Commerce

Tech Policy Press 32d ago News Agents & autonomy

Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents

Always-on agents are systems whose future behavior depends on durable state accumulated across earlier interactions. We treat them as persistent-state systems: the operative system includes retrievable memories, but also task ledgers, permissions, credentials, commitments, provenance and audit records, shared state, trigger conditions, and externally committed effects linked to those records. The survey reads the literature through six diagnostic axes for each state item, authority, scope, mutab
arXiv 33d ago Research RegulationHealthcare

Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era

What eras bookend our interregnum?
Import AI 33d ago Field notes Agents & autonomy

Beatbot Sora 70, la prova del robot da piscina più smart

Beatbot Sora 70 si prende cura in modo preciso e totale della pulizia dell'acqua dalla superficie alle profondità
Wired Italia (IT) 33d ago News Agents & autonomy

Exploration and Online Transfer with Behavioral Foundation Models

Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories. For their generality over tasks, such models are sometimes called ``Behavioral Foundation Models'' (BFMs). While they have shown strong performances and improvements in recent years, the current framework and algorithms still assume that, during the transfer phase, the ag
arXiv 33d ago Research Agents & autonomy

Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies

Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challenges they are ultimately designed to handle. However, real-world evaluation is also the bottleneck for iterating on robot policies: it is costly, difficult to reproduce, and often too sparse to reliably compare nearby model variants. A straightforward proxy for performance is validation loss on expert demonstrations, but this proxy is often poorly correlated wi
arXiv 33d ago Research Agents & autonomy

Experience Graphs: The Data Foundation for Self-Improving Agents

The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We argue that long-horizon agentic tasks -- code generation, scientific discovery, hardware design -- are such a workload. These agents explore: they generate artifacts, execute tools, observe failures, branch, and repair over hundreds of steps. This search produces a structured object we call an experience graph: executable artifacts, tool outputs, rewards, sibl
arXiv 33d ago Research Agents & autonomy

STAT+: AI scientist company Edison Scientific tapped by team behind Metsera to create new biotechs

Edison Scientific and investment firm Population Health Partners are teaming up to leverage AI agents in drug discovery and development.
STAT News (health AI, headlines) 33d ago News HealthcareAgents & autonomy
← Newer Older →