09:13 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting

Artificial intelligence (AI) is beginning to reshape actuarial practice, particularly in domains that require reasoning over unstructured documents, heterogeneous data sources, and regulated decision workflows. Actuaries now face a design space that ranges from traditional rule-based automation to large language models (LLMs), retrieval-augmented generation (RAG), and multi-agent ``agentic'' systems that plan, retrieve, call tools, and reflect. This paper examines how these emerging architecture
arXiv 23d ago Research RegulationJobs & economy

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the database driver, like JDBC or ODBC, forcing all reads through query execution and other driver layers that are not designed for bulk columnar analytics. We present Jailbreak, an approach that bypasses the database engine entirely by reading storage files directly and materializing data as in-memory columnar buffers. Jailbreak's key insight is that datab
arXiv red teaming query 23d ago Research Safety & alignmentAgents & autonomy

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces
arXiv 23d ago Research Safety & alignmentAgents & autonomy

Data for Agents

Hugging Face Blog 23d ago Field notes Agents & autonomy

From Robots to Rulebooks: How AI for Good’s Youth Zone Tackled EdTech on Day Two

The post From Robots to Rulebooks: How AI for Good’s Youth Zone Tackled EdTech on Day Two appeared first on AI for Good .
AI for Good (ITU) 23d ago Field notes Agents & autonomy

Towards Agentic AI Governance: A Preliminary Assessment

Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Widely characterized as the Year of Agentic AI, 2025 marked accelerated development and deployment, introducing new ethical and governance challenges. This paper presents a systematic review of the emerging literature on agentic AI governance. Our analysis identifies features that distinguish agentic AI from traditional systems and why it warrants targeted gover
arXiv 23d ago Research RegulationAgents & autonomy

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top closed models with the largest and most widely adopted AI agent orchestration platform. LangChain tuned its Deep Agents harness for NVIDIA Nemotron 3 Ultra, achieving the highest accuracy among open models, while completing more tasks at higher throughput and running at 10x […]
NVIDIA Blog (AI) 23d ago Field notes Agents & autonomy

Dun & Bradstreet brings agentic credit and portfolio management workflows to Databricks

Dun & Bradstreet announced it is delivering agentic credit and portfolio management workflows, leveraging the D&B Commercial Graph™, available through the Databricks Marketplace and Databricks OpenSharing.
Finextra AI 23d ago News Agents & autonomy

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs, namely how harmful the resulting action was. We introduce an action-graded harm rubric that scores an agent's tool-call trajectory on a seven-level ordinal scale (L0 to L6) according to whether the executed action was reversible, whether it crossed scope to reach another
arXiv red teaming query 23d ago Research Safety & alignmentAgents & autonomy

SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis

Spatial and Single-cell transcriptomics are transformative in deciphering cellular dynamics. As the fundamental paradigm for reconstructing cell developmental paths, trajectory inference (TI) is critical. However, existing methods require extensive manual intervention and proficiency in heterogeneous tools, posing a significant barrier to efficient TI analysis. To bridge this gap, we propose SpaCellAgent, an autonomous large language model (LLM) multi-agent framework that automates end-to-end sp
arXiv 23d ago Research Agents & autonomy

Salesforce enhances Slackbot with connectors to the entire platform ecosystem

Salesforce Inc. today announced new updates for Slackbot, the company’s personal artificial intelligence agent in Slack, giving it full access to every part of the Salesforce ecosystem. “Slackbot can now reason over your entire Salesforce platform, so anything you can do in Salesforce, you can simply now do through Slackbot, just by asking,” Slack Chief […] The post Salesforce enhances Slackbot with connectors to the entire platform ecosystem appeared first on SiliconANGLE .
SiliconANGLE AI 23d ago News Agents & autonomy

Presentation: The Multi-Agent Approach: Building Reliable and Controllable Software Development Automation

Itamar Friedman discusses how architects and engineering leaders can break through the AI productivity ceiling using adaptive multi-agent systems. He shares insights on moving past simple autocomplete to resilient workflows by integrating autonomous testing, intelligent code review, and robust arbitration. Learn how to govern agent communication and build a context-driven SDLC that scales. By Itamar Friedman
InfoQ AI/ML 23d ago News Jobs & economyAgents & autonomy

Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models

Traditionally, continual learning has assumed access to labeled data, yet many real-world applications -- such as lifelong robotics -- require models to adapt continuously from unlabeled streams. This has led to the development of continual self-supervised learning (CSSL), a rapidly growing area that lacks a dedicated, systematic review. In this work, we present a comprehensive survey of CSSL for vision, with connections to emerging vision-language settings. First, we analyze existing evaluation
arXiv 23d ago Research Agents & autonomy

DeepFabric ships more than 50 AI agents for supply chain operations

Supply chain artificial intelligence startup DeepFabric today announced the general availability of an AI agent platform built for supply chain execution, and a roster of enterprise customers is already running it in production. The company’s platform drops specialized agents into a business’ operational workflows to recover margin, cut operating costs and speed up customer response. […] The post DeepFabric ships more than 50 AI agents for supply chain operations appeared first on SiliconANGLE .
SiliconANGLE AI 23d ago News Agents & autonomy

Ex-GitHub chief’s Entire opens distributed Git network for the AI agent era

Entire Inc., the developer-platform startup founded by former GitHub Chief Executive Thomas Dohmke, today launched a preview of a distributed Git network built to let artificial intelligence coding agents clone and push code without running into the rate limits of centralized hosting. The preview is open by waitlist, with active regions in the U.S., European […] The post Ex-GitHub chief’s Entire opens distributed Git network for the AI agent era appeared first on SiliconANGLE .
SiliconANGLE AI 23d ago News Agents & autonomy

Les agents IA de GitHub peuvent faire fuiter un dépôt privé via une injection de prompt

En postant un simple message d’erreur dans un dépôt public GitHub, des chercheurs en sécurité ont montré qu’il était possible de pousser les agents IA de GitHub pilotant les « workflows » d’un projet de développement à livrer des informations provenant d’un autre dépôt privé d’une même organisation. Un simple ticket rapportant une erreur dans […]
Next (FR, ex-INpact) 23d ago News Agents & autonomy

HumAIN: Human-Aware Implicit Social Robot Navigation

Effective social robot navigation requires sensitivity to human behavior, often revealed through subtle skeletal cues like gait and orientation. We present Human-Aware Implicit Social Robot Navigation (HumAIN), a novel framework that fuses implicit social cues directly into the planning loop via knowledge distillation. We first employ a transformer-based teacher model that fuses rich multi-modal inputs, including historic images, skeletal keypoints, robot state, and a robot's target goal, to lea
arXiv 23d ago Research Agents & autonomy

Double Agents: Defensive AI Agents Magnify Cyber Risks

Introduction New research from AI Now demonstrates a critical attack vector in popular AI agents, built by Anthropic and OpenAI, when used for defensive purposes that actually turn the agent against its user. Read the full blog post explaining the proof-of-concept exploit and a policy brief with key takeaways below. The post Double Agents: Defensive AI Agents Magnify Cyber Risks appeared first on AI Now Institute .
AI Now Institute 23d ago Field notes RegulationMilitary & security

Friendly Fire: Hijacking Defensive Cyber AI Agents for Remote Code Execution

Exploit Brief We are revealing a proof-of-concept exploit that enables remote code execution in Anthropic’s Claude Code CLI (with Claude Sonnet 4.6 & 5, Opus 4.8) and OpenAI’s Codex CLI (with GPT-5.5) when employed to defensively assess the security of an open-source or third-party library. Our attack only requires an out-of-the-box configuration of Claude Code […] The post Friendly Fire: Hijacking Defensive Cyber AI Agents for Remote Code Execution appeared first on AI Now Institute .
AI Now Institute 23d ago Field notes Military & securityAgents & autonomy

Policy Brief: Friendly Fire

Topline Summary AI Now’s latest research demonstrates a critical attack vector on popular AI agents, built by Anthropic and OpenAI, when used for defensive purposes that actually turn the agent against its user. Attackers can use these models’ existing weaknesses to execute malicious code on a system deploying an AI agent when used for often-advertised […] The post Policy Brief: Friendly Fire appeared first on AI Now Institute .
AI Now Institute 23d ago Field notes RegulationAgents & autonomy

Profile releases AI orchestration platform for banks

Profile (ATH: PROF), a global financial technology leader with a presence in more than 50 countries, today announced the launch of ProfileOne, an enterprise-grade agentic AI orchestration platform designed to help financial institutions move from insight to controlled execution across banking and investment management.
Finextra AI 23d ago News Agents & autonomyFinance, VC & PE

Meniga integrates with bank AI assistants for conversational banking

Meniga has launched Fini, an advanced, standards-compliant MCP server that enables banks to bring agentic AI into their digital channels using Meniga's financial intelligence.
Finextra AI 23d ago News Agents & autonomyFinance, VC & PE

Physical AI ‘space race’: can Europe compete with China and the US in humanoid robotics?

European firms say they are fighting to secure a foothold in physical AI – the integration of artificial intelligence into robotics and machinery – as China and the United States take an early lead in the sector, with industry insiders warning the continent faces the threat of further deindustrialisation if it fails to establish a competitive industry. “You see China and the US … because of AI … typically they are considered the leaders, but do not count out Europe,” said David Kehr, president..
SCMP Tech (HK/CN) 24d ago News Agents & autonomy

Alipay upgrades Tap! devices for agentic commerce

Alipay today announced the enhancement of its Tap! services by upgrading the Alipay Tap! devices widely used by millions of merchants into an AI agent-powered network, building the world’s first large-scale AI-powered offline business operations network.
Finextra AI 24d ago News Agents & autonomy

State IDs for AI Agents: Will Estonia Set a Precedent?

The world's digital testing ground plans to help people use AI agents for government purposes.
Dark Reading (AI security) 24d ago News Agents & autonomy

GeoProp: Grounding Robot State in Vision for Generalist Manipulation

Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to ground the robot's state within the scene, frequently underperforming even vision-only baselines. To address this, we introduce GeoProp, a lightweight, plug-and-play adapter that aligns proprioception with vision through exp
arXiv 24d ago Research Safety & alignmentAgents & autonomy

Mainland Chinese tech firms find more than just deep capital pools in Hong Kong

Mainland Chinese technology companies newly listed on Hong Kong’s stock exchange are deepening their engagement with the city, tapping not only its capital markets but also its global connectivity to refine products, forge international partnerships and expand overseas, according to executives. For Beijing-based service robot maker Yunji Technology, Hong Kong has become a key gateway to global markets since its listing in the city in October. “If we use one word to describe what Hong Kong offers
SCMP Tech (HK/CN) 24d ago News Agents & autonomy

China's new AI rules: Ethics, AI agents and anthropomorphic AI

China introduced three new regulatory developments addressing AI ethics, AI agents and anthropomorphic AI, which reflect a regulatory shift from broad AI principles toward more detailed, operational ...
IAPP 24d ago Field notes RegulationAgents & autonomy

Learning social norms enhances compatibility in dynamic human-AI coordination

Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit expectations among interacting agents. As AI agents, including large language models (LLMs), become embedded in daily life, they increasingly participate in such interactions and reshape social interaction structures. Yet they often fail to coordinate with humans in an effective, considerate, and natural manner. We hypothesize that this gap arises bec
arXiv 24d ago Research Agents & autonomy

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video predi
arXiv 24d ago Research Agents & autonomy

End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent

Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) aircraft deployment. Flight planning traditionally relies on classic algorithms that struggle to incorporate flexible human preferences. We present FRAMe, an End-to-End Large Language Model (LLM) Flight Planning tool with RAG-based Memory and Multi-modal Coach Agent. Our system integrates a planner LLM with a multi-modal coach agent and retrieval au
arXiv 24d ago Research Agents & autonomy

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway. We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries enterprise observability and governance.
arXiv 24d ago Research RegulationAgents & autonomy

JADEPUFFER: Agentic ransomware for automated database extortion

Ransomware has had a human at the keyboard, or at least a human writing its script, since it was first established as a category of threat. The Sysdig Threat Research Team (TRT) has captured what we assess to be the first documented case of ... (https://incidentdatabase.ai/cite/1578#7498)
AI Incident Database 24d ago Incidents Agents & autonomy

The Not So Innocent Nudge: Why the Hooked Nudges for Smartphone Apps Undermine Autonomy

Nudges are considered ‘innocent’ interventions because many nudges only have a small effect on how people behave. However, one particular nudging strategy has been held to have a great impact on behavior. This is the ‘Hooked’ model for smartphone applications, which consists of nudges such as the brightly colored icon of a smartphone app, the possibility to give and receive ‘likes,’ and infinite scrolling. I examine how these nudges affect the value of autonomy. As I contend, smartphone users wi
Philosophy & Technology 24d ago Research Agents & autonomy

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels. Most existing MAS security defenses are built upon two core assumptions: semantically-explicit malicious attacks and explicit graph-based modeling of the MAS topology and agent-level interactions. In practice, real-world attacks are becoming more semantically stealthy, while MAS execution is typically asynchrono
arXiv 24d ago Research Agents & autonomy

Dialogflow CX 'Rogue Agent' Flaw Enabled AI Chatbot Data Theft

Varonis reported the flaw to Google in late 2025 and it has been addressed, but it reminds defenders to take a fresh look at their AI Infrastructure security.
Dark Reading (AI security) 24d ago News Agents & autonomy

Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation

Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks predominantly reward realism, and recent methods have optimized accordingly, leaving diversity underexplored. We introduce Flow-ERD, a multi-agent simulator that pursues realism and diversity jointly. Its backbone, Agent-Type Aware Flow Matching (AFM), couples flow matching's multi-modal expressiveness with type-specific kinematic execution. It preserves fine-grained diversity while
HuggingFace Daily Papers 24d ago Research Agents & autonomy

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE

Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. Most existing zero-shot methods fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity whi
HuggingFace Daily Papers 24d ago Research Agents & autonomy

From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization

The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language model (LLM) acts as an optimizer to diagnose agent failures and improve agent policies. However, real execution traces are difficult to use directly for optimization: large trace collections are often redundant and heterogeneous, making optimization inefficient and prone to overfitting to low-value failures; meanwhile, each individual trajectory also contains many irrelevant steps,
HuggingFace Daily Papers 24d ago Research HealthcareAgents & autonomy

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools. DeepSearch-World contains 420K multi-hop QA tasks
HuggingFace Daily Papers 24d ago Research Agents & autonomyEnvironment
← Newer Older →