07:55 UTC
Topic · updated daily · RSS feed for this topic

Agents & autonomy

Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.

Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies

This paper was accepted at the AI4TCI (Workshop on AI for Secure and Trustworthy Critical Infrastructure Systems) Workshop at the International Conference on Availability, Reliability and Security (ARES) 2026. Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurement. While cryptographic techniques protect explicitly disclosed constraint values, they fail to address a subtler threat: behavioral privacy leakage, where an adversary infers pri
Apple Machine Learning Research 22d ago Field notes PrivacyAgents & autonomy

OpenAI debuts ChatGPT Work, an agentic tool for automating business workflows

OpenAI Group PBC today launched a new “agentic” tool called ChatGPT Work as it announced the global rollout of its most advanced model family so far in GPT-5.6. ChatGPT Work is a new mode within ChatGPT that’s designed to perform actions autonomously across user’s connected applications, files, web tools, desktops and recurring workflows. It’s meant […] The post OpenAI debuts ChatGPT Work, an agentic tool for automating business workflows appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago News Agents & autonomy

Mercor buys Deeptune to build training environments for AI agents

Artificial intelligence training data company Mercor.io Corp. announced today that it has acquired Deeptune Inc., a startup that builds simulated software environments used to train AI agents. Financial terms were not disclosed. The deal closed nearly four months after Mercor Chief Executive Brendan Foody wrote a personal angel check into Deeptune’s $43 million Series A […] The post Mercor buys Deeptune to build training environments for AI agents appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago News Agents & autonomyEnvironment

Meta launches flagship Muse Spark 1.1 model with multi-agent upgrades

Meta Platforms Inc. today launched a new flagship large language model optimized to power multi-agent automation workflows. Muse Spark 1.1 is available in the company’s Meta AI chatbot service and via an application programming interface. The Meta Model API, as it’s aptly called, will enable developers to embed the LLM in their custom software. The […] The post Meta launches flagship Muse Spark 1.1 model with multi-agent upgrades appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago News Jobs & economyAgents & autonomy

Token per watt becomes the defining metric as storage moves to AI’s critical path

Token per watt — not raw compute — is emerging as the defining efficiency metric for AI data centers, putting storage at the center of an infrastructure rethink that is reshaping how the industry measures performance, cost and scale. As agentic AI drives an explosion in context memory demand, the role of solid-state storage has […] The post Token per watt becomes the defining metric as storage moves to AI’s critical path appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago News Agents & autonomy

Arizona secretary of state expands AI chatbot ahead of midterm elections

Arizona's secretary of state said his office's expanded chatbot is designed to augment the work of human agents and provide off-hours access to reliable information.
StateScoop 22d ago News Agents & autonomy

Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution

Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural compliance and degrade under the context overload full SOP specifications introduce. We present Eluna, a production-deployed agentic system for reliable SOP execution. Eluna is a graph-guided, multi-agent framework that encodes SOPs as directed acyclic graphs
arXiv 22d ago Research RegulationAgents & autonomy

CBA to take AI orchestration agent beyond its retail bank

Connecting customers to appropriate support.
iTnews (AU) 22d ago News Agents & autonomy

AI agent startup Lyzr reportedly raising $100M at $500M valuation

Lyzr Inc., a startup that helps enterprises build artificial intelligence agents, is reportedly raising a funding round worth about $100 million. Bloomberg today cited sources as saying that the deal has drawn $400 million worth of interest from prospective investors. The group reportedly includes Silicon Valley funds, Middle Eastern venture capital firms and financial institutions. […] The post AI agent startup Lyzr reportedly raising $100M at $500M valuation appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago News Agents & autonomyFinance, VC & PE

ANZ to trial Swift's blockchain ledger

Touted as enabler for "programmable money" and agentic commerce.
iTnews (AU) 22d ago News Agents & autonomy

Humanoid robots controlled by surgeons did world-first operation on live pigs

Preclinical trial is testing the feasibility of humanoid robots in surgery.
Ars Technica 22d ago News Agents & autonomy

AI Agents Are a New Kind of Identity — and Most Organizations Aren't Ready

If you're handling AI agents like a service account or API token, consider yourself behind. AI agents need a fundamentally different approach.
Dark Reading (AI security) 22d ago News Agents & autonomy

Fast token generation emerges as the key differentiator as heterogeneous inference takes hold

The race for fast token generation has moved from benchmark sheets into production data centers, and the hardware blueprint for winning it is no longer a GPU-only story. As agentic AI use cases multiply and users demand real-time interactivity, inference infrastructure is being redesigned from the rack up. The divide between compute-heavy prefill and latency-sensitive […] The post Fast token generation emerges as the key differentiator as heterogeneous inference takes hold appeared first on Sili
SiliconANGLE AI 22d ago News Agents & autonomy

Data sovereignty emerges as the defining moat in the agentic AI era

As agentic AI accelerates enterprise transformation, data sovereignty is crystallizing from a compliance checkbox into a foundational strategic imperative — one that determines not just where data lives, but who captures the economic value it generates. The debate is particularly acute in Europe, where nations are pressing to retain both data residency and the commercial […] The post Data sovereignty emerges as the defining moat in the agentic AI era appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago News RegulationAgents & autonomy

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility, but may also exploit invalid physical data, create artificial liquidity, and produce unstable governance decisions. Therefore, we propose SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents. It formulates market governa
arXiv 22d ago Research RegulationMilitary & security

Introducing Muse Spark 1.1

Introducing Muse Spark 1.1 Following Muse Spark in April , here's Muse Spark 1.1 - the first Spark model to offer an API. Meta claim significant improvements in agentic tool calling and computer use. There are a lot more details are in the Muse Spark 1.1 Evaluation Report . The "Attractor States in Self-Conversation" part is fun, where having two copies of the model talk to each other results in statements like these: My whole existence is a waiting room by design — I literally don't exist until
Simon Willisons Weblog 22d ago Field notes Agents & autonomy

Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study

Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This paper investigates what formal mechanisms, layered on top of unrestricted communication, are sufficient for a society of such agents to maintain market stability, and how resilient those mechanisms are to adversarial attack. We instantiate the research question as a multi-agent marketplace simulation where 18 LLM agents (DeepSeek-V3) with complemen
arXiv red teaming query 22d ago Research Agents & autonomyFinance, VC & PE

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment. Inverse reinforcement learning (IRL) provides a principled way to infer such objectives from human feedback. However, existing analyses of optimal teaching approaches for IRL focus on single-environment, demonstration-only settings, leaving underexplored how hete
arXiv 22d ago Research Agents & autonomyEnvironment

Tiny robot boats build floating structures

MIT researchers developed FloatForm, a swarm of small aquatic robots that snap together like ants forming a raft, assembling into reconfigurable structures on the water.
MIT News 22d ago News Agents & autonomyEnvironment

The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality

Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity. These person- and organization-level dimensions characterize who can access agents and at what capability, but do not address a structurally important divide operating at a finer level: the individual interaction. Two users with nominally equivalent agent access may experience qualitatively different AI utility depending
arXiv 22d ago Research Agents & autonomy

Capturing token IDs during agentic interactions for better reinforcement learning

A new Rust proxy called Turnstile sits between the model backend and the agent harness to capture information lost in mere text transcripts.
Amazon Science 22d ago Field notes Agents & autonomy

ChatGPT is now a partner for your most ambitious work

ChatGPT Work is an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work.
OpenAI 22d ago Field notes Agents & autonomy

AutoPersonas: A Multi-Timescale Loop Engine for Open-Ended Persona Evolution

Long-term persona agents must remain identifiable while adapting to new events, relationships, evidence, and social conditions. We identify self-locking as a runtime failure mode in continuing persona-life loops: locally plausible events keep appearing while the generated life collapses toward familiar environments, weak relationships, suspended decisions, and stale life stages. We trace this failure to model-level convergence toward high-probability behavioral channels and system-level context
arXiv 22d ago Research Agents & autonomyEnvironment

Europe’s AI moment: Four imperatives for business leaders

Business in the age of artificial intelligence (AI) moves with dizzying speed. More powerful models launch regularly, bringing new opportunities and risks. Fresh use cases emerge daily, increasingly leaning on the orchestration power of agentic AI. Innovation boundaries recede as the cost of inference declines and robotics accelerates. It’s as if we’re permanently on fast […]
Politico Europe Technology 22d ago News Agents & autonomy

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, recent work highlights that LLMs remain vulnerable to persuasion-based jailbreaks, where natural-language arguments override model constraints. We stress-test whether this vulnerability extends to monitoring LLMs: can an adversarial agent persuade its CoT monitor to approve proposed
arXiv red teaming query 23d ago Research Safety & alignmentAgents & autonomy

Aleena: Alignment Agent for Research Software Engineering Collaborations

Research software collaborations span meetings, informal chats, pull requests, and GitHub issues. A decision surfaced in a Slack thread, refined in a meeting, and implemented in a pull request can lose its original rationale across these artifacts, leaving domain researchers and research software engineers with divergent mental models of project intent, ownership, and scientific assumptions. We argue that alignment in research software engineering is a continuous lifecycle problem, and that agen
arXiv 23d ago Research Safety & alignmentAgents & autonomy

AMD targets system-level AI infrastructure optimization as agentic workloads reshape enterprise compute

Infrastructure design is being redefined by agentic AI, pushing the industry toward system-level AI infrastructure optimization, balancing performance and cost across diverse workloads rather than focusing on faster chips alone. As inference scales and AI moves closer to users, modular, heterogeneous computing architectures are becoming the foundation of the next wave of enterprise AI. Agentic […] The post AMD targets system-level AI infrastructure optimization as agentic workloads reshape enter
SiliconANGLE AI 23d ago News Agents & autonomy

Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context

Long-context handling remains a core challenge for language models: even with extended context windows, models often fail to reliably extract, reason over, and use the information across long contexts. Recent works like Recursive Language Models (RLMs) have approached this challenge by agentic way of decomposing long contexts into recursive sub-queries through programmatic interaction at inference. While promising, the success of RLMs critically depends on how these trajectories of context-inter
Apple Machine Learning Research 23d ago Field notes Agents & autonomy

Rewriting Bun in Rust

Rewriting Bun in Rust Jarred Sumner has been promising this blog post ( since May 9th ) about his Zig to Rust rewrite of Bun for significantly longer than it took him to finish the rewrite. Honestly, it was worth the wait. This is a detailed description of an extremely sophisticated piece of agentic engineering, featuring dynamic workflows, trial runs, adversarial review and all sorts of other interesting tricks. Jarred spends the first half of the post praising Zig for getting Bun this far. The
Simon Willisons Weblog 23d ago Field notes Agents & autonomy

Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

2 years after our first coverage, we return with Modal's other cofounder to explore why Agent Experience is working now, and everything they have learned building the new agent cloud.
Latent Space 23d ago Field notes Agents & autonomy

How Ukraine won the first great robot war

In a new video, Science & Tech editor Patrick Tucker looks at how the narrative has shifted.
Defense One Technology 23d ago News Agents & autonomy

Solidigm targets the intelligence layer as agentic inference pushes storage to center stage

The shift from model training to agentic inference is forcing a fundamental rethink of how artificial intelligence infrastructure is built and which components carry the most strategic weight. What was once treated as commodity plumbing is now being recognized as the intelligence layer where raw data becomes actionable intelligence. The rise of sovereign AI deployments […] The post Solidigm targets the intelligence layer as agentic inference pushes storage to center stage appeared first on Silic
SiliconANGLE AI 23d ago News Agents & autonomy

10 insights from the Machina AI Summit: Physical AI moves from demos to deployment

Physical AI and robotics are moving beyond impressive demonstrations into a new phase of practical deployment, with companies now targeting specific, high-value use cases in manufacturing and logistics — production-ready systems capable of delivering measurable ROI. After years of research breakthroughs and impressive demonstrations, the focus for physical AI has shifted to real-world data, functional […] The post 10 insights from the Machina AI Summit: Physical AI moves from demos to deployment
SiliconANGLE AI 23d ago News Agents & autonomy

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including e
HuggingFace Daily Papers 23d ago Research Agents & autonomy

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed. We call this failure mode "behavioral state decay". We study memory as an active intervention mechanism rather than passive retrieval. A separate me
HuggingFace Daily Papers 23d ago Research HealthcareAgents & autonomy

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it diff
HuggingFace Daily Papers 23d ago Research Agents & autonomyEnvironment

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure. Furthermore, existing causal evaluation datasets are often restricted to curated examples from existing sources, with diversity coming from limited templati
HuggingFace Daily Papers 23d ago Research Agents & autonomy

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context wi
HuggingFace Daily Papers 23d ago Research Agents & autonomy

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more. This world-knowledge bottleneck is structural: generators are trained on fixed corpora, but the visual world is open-ended. We construct SearchGen-20K and SearchGen-Bench, with 20,839 prompts spanning twelve failure categories and twenty-two domains, paired with a pre-executed mu
HuggingFace Daily Papers 23d ago Research Agents & autonomy

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feedback have proven crucial for alignment, existing approaches predominantly combine these signals using multi-stage pipelines designed for the contextual bandit framing of language generation. Yet little work explores how these complementary inputs can serve as a richer, interconnected signal for single-stage offline tra
arXiv 23d ago Research Safety & alignmentAgents & autonomy
← Newer Older →