Topic · updated daily · RSS feed for this topic
Agents & autonomy
Agentic AI acting in the world: oversight, incidents, robotics and the governance questions agents raise, daily.
Google Bets 'Agentic Defense' Strategy Can Outpace Attackers
Google Cloud incorporates key Wiz capabilities into an agentic defense platform to automate threat detection and remediation against AI attacks.
Agentic Synthesis against Counterexample-Supplemented Sketches
Coding agents can fix a failing example without preserving the domain rule that made it fail, so later generations can repeat the same plausible mistake. We present agentic synthesis against counterexample-supplemented sketches, a repository-native method for systems whose governing policy is discovered during implementation. A human starts with a partial, code-shaped sketch, and a coding agent generates the first implementation. When a concrete failure exposes missing or mistaken policy, an ope
Knowledge-Centric Agents for Workflow Generation
Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON generation task, struggling with structural brittleness and lacking the experiential knowledge required for effective design. We argue that successful workflow generation requires modeling knowledge itself, including its structure, hierarchy, and reason
New coding, robotics lab to empower 650 learners
Isuzu Motors sets up a R2 million AI, coding and robotics lab at Khulile Primary School in the Eastern Cape.
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult. Existing evaluators use different rubrics and evidence sources and may fail on JavaScript-rendered pages or repository-specific identifiers. For 50 datasets from 10 repositories, the standard deviation of normalized scores across available tools averages 15.0 percentage points and reaches 30.3 for one dataset. Because these outputs are not equivalent measur
A Humanoid Company Backed by Eric Trump Is Preparing Its Robots for War
The CEO of Foundation Future Industries, which counts the president’s son as its chief strategy adviser, tells WIRED it’s exploring some “kinetic things.”
QCon AI Boston: Production AI Moves Beyond Prompts to Platforms, Harnesses, and Evals
QCon AI Boston 2026 focused on the operational challenges of deploying AI agents, emphasizing the need for robust production infrastructure. Key themes included improving context management, ensuring security through a "harness" around agents, and adopting a comprehensive engineering model for AI. By Tatiana Fesenko
Making Agent-Mediated Contributions Governable: A Project-Level Governance Manifest for Open-Source AI Collaboration
Generative AI and coding agents are intensifying a central governance tension in open-source software (OSS): they scale contribution generation faster than maintainers can assess risk, evidence, and accountability. Existing responses improve agent-readability and traceability, but project rules must also organize contribution-specific risk, evidence, accountability, and review-gate states. We theorize this organizational arrangement as project-side governability infrastructure. A diagnostic audi
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understand
Don’t block the bots. Build the gate
Media companies have gotten pretty good at blocking AI bots. They have been less good at deciding which ones should be allowed in. Reports say most major publications almost universally block crawlers from major AI companies like OpenAI and Anthropic, and many small to midsize publishers mirror that, configuring their robots exclusion protocol settings, or robots.txt, to keep them out. It’s an understandable stance, and it might even be appropriate. But it also crudely compresses a complex issue
KI in Unternehmen: Rendite steigt, Kontrolle hinkt hinterher
Deutsche Unternehmen erwarten bis 2026 einen ROI von 24 Prozent durch KI. Doch viele sind noch nicht bereit für den Einsatz von KI-Agenten.
Broken Drone, Far from Home: The Case for Overseas Autonomous System Sustainment
It’s 2028, and the Fujian carrier strike group has just left Yulin, China, for an unknown destination. The U.S. Navy’s “Hedge Strategy” has let commanders disperse unmanned systems across regional choke points as forward-deployed scouts. But days before contact with Fujian, an unmanned undersea vehicle in the Banda Sea transmits an error code and must head back to Yokosuka for repairs — a 3,000-mile, 15-day trip. A second platform, an unmanned surface vehicle, is overdue for a routine oil filter
China’s BrainCo unveils mind-to-robot platform at World AI Conference
Chinese tech unicorn BrainCo has unveiled what it claims is the world’s first integrated “brain-to-robot” platform that lets users control robots using only their thoughts, without moving a muscle. The launch of the Brain-Controlled Robot AI Platform was announced on Friday at the World Artificial Intelligence Conference (WAIC) in Shanghai – China’s premier AI event. It comes as global technology firms are racing to develop powerful embodied AI systems, where artificial intelligence is...
OpenAI: Erstes Hardware-Gerät ist ein Makro-Pad für 230 US-Dollar
OpenAIs erstes eigenes Gerät heißt Codex Micro und ist ein Makro-Pad. Es kostet 230 US-Dollar und steuert KI-Agenten, wird aber nicht von OpenAI selbst gebaut.
Apple Upgraded to Buy by HSBC on Agentic AI, Hardware Pipeline
Apple Inc. was upgraded to buy at HSBC Holdings Plc on Friday, in the latest reflection of how Wall Street is increasingly positive on the iPhone maker’s position in a market where many trades related ...
The labor-saving paradox in transition: An empirical study on the inverted U-shaped relationship between industrial robot application and overtime work in Chinese firms
Publication date: November 2026 Source: Technological Forecasting and Social Change, Volume 232 Author(s): Man Qin, Yaoyao Zhang
US-Robotikforscherin vom MIT erhält bayerischen High-Tech-Preis
Die US-Amerikanerin Daniela Rus erhät den bayerischen High-Tech-Preis des Ministerpräsidenten von Bayern. Die Forscherin hat besondere Beziehungen zu Bayern.
Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving
Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the model on alternate simulation ticks and replaying the previous command in between, so half of all control outputs ignore the newest observations. We present a fast-slow architecture that removes this compromise. A frozen 7B vision-language backbone acts as the slow syst
Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers
arXiv:2607.14447v1 Announce Type: new Abstract: AI assistants increasingly retrieve web content at inference time to provide fresh and grounded answers, yet it remains unclear whether these search-augmented capabilities respect website-owner restrictions expressed through robots.txt. We present a controlled empirical study of ten widely used AI assistants with advertised web-search capabilities. For each assistant, we first identify a configuration that actually produces observable web-browsing
Traccia: An OpenTelemetry-Based Governance Platform for AI Systems
arXiv:2607.14309v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) and Artificial Intelligent (AI) powered autonomous agents has fundamentally changed the existing forms of software governance. In spite of the rigorous standards of transparency and account ability required according to the international frameworks such as the European Union's AI Act, there is a considerable gap between theory and reality. The present study discusses the inherent drawbacks of
PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction
Equipping humanoid robots with coherent and adaptable personas is crucial for fostering natural, engaging, and trustworthy human-robot interaction (HRI). However, existing approaches often rely on static, hard-coded identities that lack the flexibility to adapt to individual user contexts. In this paper, we present PACE (Persona Adaptation through Conversational Elicitation), a novel framework for the interactive generation and deployment of structured personas on the Ameca humanoid robot. Our s
PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction
Equipping humanoid robots with coherent and adaptable personas is crucial for fostering natural, engaging, and trustworthy human-robot interaction (HRI). However, existing approaches often rely on static, hard-coded identities that lack the flexibility to adapt to individual user contexts. In this paper, we present PACE (Persona Adaptation through Conversational Elicitation), a novel framework for the interactive generation and deployment of structured personas on the Ameca humanoid robot. Our s
From humanoids to Huawei: what to watch as Xi attends China’s WAIC amid US AI rivalry
China will use its largest annual artificial intelligence gathering this week to showcase an ambition that extends beyond catching up with the United States in foundation models to building dominance in autonomous agents, scientific research, humanoid robots and consumer devices. The World AI Conference (WAIC) in Shanghai comes amid an intensifying technology rivalry with the US, whose restrictions continue to constrain China’s access to advanced computing chips. Beijing has responded by making.
A multimodal dataset for socially aware navigation of heavy-duty construction robots
Blue Water Autonomy, Saildrone launch lawsuits against Navy over MUSV Marketplace
Both companies claim that their proposals did satisfy Navy requirements for the new MUSV marketplace.
Agentic AI Is Untamable: Ask the Right Security Questions
Forget about attackers. Agentic artificial intelligence is creating enough risks for organizations and demands a security reframe.
Volkswagen enters the robotaxi race with a shared shuttle service in Hamburg
Volkswagen’s autonomous mobility subsidiary Moia has begun offering rides in self-driving ID Buzz vans to preregistered residents in Hamburg, marking the first time a major European automaker has launched an autonomous passenger service on its home continent. Up to five vehicles are operating at initial launch, with the fleet expected to expand to 10, and […] This story continues at The Next Web
Fear of humanoid robots spurs human workers to strike at Hyundai auto factory
Hyundai aims to deploy 25,000 Atlas robots starting with US factories in 2028.
DSWorld: A Data Science World Model for Efficient Autonomous Agents
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects of data science operations before real execution. In this paper, we introduce the concept of Data Science World Model, which model the data science execution environment by predicting environment state transitions conditioned on current workflow sta
Recursive Harness Self-Improvement
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a ta
RecGPT-V3 Technical Report
Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior mo
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloa
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 in
When Does Muon Help Agentic Reinforcement Learning?
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneous action can lead to irreversible consequences. Existing safety mechanisms are primarily reactive, lacking the ability to assess risks before execution. In this paper, we introduce SeerGuard, a consequence-aware safety framework designed to mitigate these risks through pre-execution instruction-level screening and acti
Un agent chargé du prompteur de Donald Trump, suspecté d’avoir parié de l’argent sur ses discours, a été suspendu
Il utilisait la plateforme de prédiction Kalshi, qui permet notamment aux internautes de parier sur la possibilité qu’une phrase ou un mot soient prononcés, et avait ainsi encaissé plus de 100 000 dollars.
The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity, and most agents still share credentials; and only three in ten isolate their highest-risk agents. The security stack is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agen
An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions
Large language models (LLMs) are increasingly embedded in clinical and population health workflows, including conversational agents such as health chatbots. As chatbots evolve from rule-based approaches to hybrid and LLM-enabled designs, risks and concerns about deployment readiness shift. Unlike rule-based chatbots, LLM outputs can be unpredictable, error-prone, and difficult to validate with traditional evaluation methods. Public health teams integrating customized LLMs into interventions face
RoboTTT: Context Scaling for Robot Policies
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturb
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 in