23:21 UTC

Research (187)

Analysis of hotspot areas in China's satellite internet innovation policies and research on policy evolution

Publication date: October 2026 Source: Telecommunications Policy, Volume 50, Issue 9 Author(s): Shuyu Pan, Ye Yuan, Wenle Jiang, Zixin Xu, Zhelun Zhu, Jiacheng Liu
Telecommunications Policy 18h ago Regulation

Medical robotics beyond automation: Human-robot collaboration and the RONNA system as a socio-technical case study

Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Marina Raguž, Domagoj Dlaka, Marko Švaco, Petar Marčinković, Dominik Romić, Filip Šuligoj, Bojan Šekoranja, Darko Chudy, Bojan Jerbić
Technology in Society 18h ago Jobs & economyHealthcare

The role of technology and exports in shaping skill- and gender-differentiated employment in global value chains

Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Mohd Shuaib, Mohammad Haseeb, Fei Fan
Technology in Society 18h ago Jobs & economy

Algorithmic transparency and citizen trust in digital governance: A cross-national analysis of AI adoption in public services

Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Ye Zheng, Muhammad Farhan
Technology in Society 18h ago RegulationTransparency

A scoping review of generative AI-powered agentic AI in education: Research landscape, agentic capabilities, and insights from the frontier agent paradigm, exemplified by OpenClaw

Publication date: Available online 28 July 2026 Source: Computers and Education: Artificial Intelligence Author(s): Ningxia Wang, Di Zou, Haoran Xie, S.Joe Qin
Computers and Education: Artificial Intelligence 18h ago Children & educationAgents & autonomy

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

arXiv:2607.26062v1 Announce Type: new Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intellectual disabilities (ID). Objective: The study aims to identify and measure representational differences related to people with ID and examine them to identify implicit biases inherent in AI chat generation technologies. Methods: Utilizing the GPT-4-Turbo model, we requested story-generation based on
arXiv cs.CY 19h ago Bias & fairnessFinance, VC & PE

Archetypes or ability? Clustering for modelling student mathematical competence

arXiv:2607.26063v1 Announce Type: new Abstract: Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different strengths. In this work, we explore the validity of these assumptions by applying clustering methods to a large dataset of 119,034 students, spanning 13 national-level exams sat in the United Kingdom and collected by the platform.
arXiv cs.CY 19h ago Children & education

The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

arXiv:2607.26064v1 Announce Type: new Abstract: AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversight. As seen by increased submissions to ML venues, the verification gap between scientific output and our ability to check it is already widening, and autonomous agents make it worse by magnitudes given human-agent asymmetry. We argue that science must evolve its verification infrastructure, as it ha
arXiv cs.CY 19h ago Agents & autonomy

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

arXiv:2607.26067v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic
arXiv cs.CY 19h ago Safety & alignmentChildren & education

AI Security Priorities: A Field-Wide Agenda

arXiv:2607.26069v1 Announce Type: new Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen. This paper presents a prioritized agenda for advancing AI security, informed by structured interviews with leaders across industry, government, and civil society, and refined through a multi-sector expert workshop. Participants identified and ranked the highest-importan
arXiv cs.CY 19h ago Military & security

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

arXiv:2607.26317v1 Announce Type: new Abstract: Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natur
arXiv cs.CY 19h ago Healthcare

"Nobody Did This": Contribution, Originality, and Accountability in Agent-Mediated Collaboration

arXiv:2607.26387v1 Announce Type: new Abstract: Collaborative knowledge work is changing in ways that go beyond disclosure or transparency. LLM agents are now embedded in how teams research, design, write, and decide: mediating between members, synthesizing inputs, reformulating ideas, and drafting shared outputs. They do not only facilitate collaboration; they operate within the workflow at the moment contributions are being formed. In doing so, they risk undermining the social conditions under
arXiv cs.CY 19h ago Agents & autonomyTransparency

Anticipatory Data Governance in the Age of AI: Emerging Signals in Data Access, Reuse, and Sovereignty

arXiv:2607.27029v1 Announce Type: new Abstract: This paper reports findings from a structured participatory foresight study comprising two expert forecasting studios convened by The GovLab between 2025 and 2026. The studios brought together nineteen senior practitioners spanning official statistics, digital and trade policy, open science, AI governance, geospatial systems, and public-sector innovation across multiple jurisdictions. Applying a qualitative signal-scanning methodology grounded in t
arXiv cs.CY 19h ago Regulation

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize quantitative constraints on macro-socioeconomic stability. As a result, AI systems may satisfy regulatory requirements while contributing to labor displacement, rising inequality, and reduced economic resilience. We introduce the Human Utility Factor (HUF), a differentiable welfare metric that mode
arXiv cs.CY 19h ago RegulationJobs & economy

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

arXiv:2607.26121v1 Announce Type: cross Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness. We define trustworthy embodied intelligence as the sustained capacity to execute specified tasks reliably under environmental and system variation while maintaining risk within acceptable bo
arXiv cs.CY 19h ago Environment

On Exercising Governance Power in Decentralized Autonomous Organizations

arXiv:2607.26204v1 Announce Type: cross Abstract: A decentralized autonomous organization (DAO) is a governance entity that allows its stakeholders to manage blockchain-based protocols through smart contracts. The DAO explicitly specifies how stakeholders make and enforce decisions concerning a protocol's operation in a smart contract, aptly referred to as its governance contract. The design of this governance contract, therefore, has far-reaching implications for the security (trust) and privac
arXiv cs.CY 19h ago Regulation

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

arXiv:2607.26313v1 Announce Type: cross Abstract: Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are metadata-borne: a stale price or a superseded record, perfectly well-formed in the payload and betrayed only by freshness, lineage, or provenance. Such a defect never enters the agent's context, and an agent cannot doubt data it cannot see. On a priced replenishment benchmark, a competent agent sile
arXiv cs.CY 19h ago Agents & autonomy

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

arXiv:2607.26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independe
arXiv cs.CY 19h ago Regulation

Constitutional Midtraining: Content Presence Drives Alignment Gains

arXiv:2607.26654v1 Announce Type: cross Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curricu
arXiv cs.CY 19h ago Safety & alignment

Hearsay: Vision-Language Medical Diagnoses Without an Image

arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts t
arXiv cs.CY 19h ago Healthcare

Can AI agents conduct open-ended AI research? Early evidence from two case studies

arXiv:2607.27191v1 Announce Type: cross Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D auto
arXiv cs.CY 19h ago Agents & autonomy

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

arXiv:2604.24155v4 Announce Type: replace Abstract: The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically. This paper extends that challenge by examining tw
arXiv cs.CY 19h ago Safety & alignmentAgents & autonomy

Optimal Causal Annotations: An Application to Casenotes in Social Services

arXiv:2502.10605v4 Announce Type: replace-cross Abstract: Problem definition: Estimating causal effects of interventions is central to policy and operations, but outcome data are often missing or costly to obtain. LLMs can provide text annotation at scale but may be subject to unknown bias. When ground-truth outcomes require expensive expert labeling or follow-up, budget limits typically allow only a fraction of the data to be labeled. Motivated by collaboration with a nonprofit conducting stree
arXiv cs.CY 19h ago Bias & fairnessRegulation

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

arXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and exam
arXiv cs.CY 19h ago Healthcare

The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning

arXiv:2507.04398v3 Announce Type: replace-cross Abstract: Generative AI is becoming part of academic writing, but its educational value depends on how control is shared between learner and system. This study examined an agency gap: performance differences that may arise when AI agent initiative is misaligned with learners' generative AI literacy. Seventy-nine medical and nursing students completed two multimodal analytical writing tasks using healthcare simulation data visualisations. They were
arXiv cs.CY 19h ago Safety & alignmentHealthcare

Statistical laws and linguistics differ in naturalistic video and fictional conversations

arXiv:2512.18072v3 Announce Type: replace-cross Abstract: Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations unfold in time is through statistical patterns such as Heaps' law, which holds that vocabulary size scales with document length. Little work on Heaps' law has looked at conversation and considered how language features im
arXiv cs.CY 19h ago Regulation

Feedback modalities in human-cobot collaboration: experimental evaluation of performance, user experience, and physiological responses

Collaborative robots (cobots) are increasingly deployed in industrial as well as non-industrial domains to support human-centered operation. While physical safety and task efficiency have received considerable attention, less is known about how feedback modality influences operator experience and physiological responses under different collaboration demands. This study examines the effects of feedback modalities in two human–cobot collaboration scenarios representing distinct coordination struct
Frontiers in Robotics and AI 23h ago Agents & autonomy

Object-grounded embodied picking for e-commerce warehouse fulfillment: a foveated diffusion policy for operational robustness

Embodied picking for e-commerce fulfillment remains vulnerable to dense clutter, reflective packaging, and background variation, which can undermine the effectiveness and robustness of visuomotor policies learned from demonstrations. A key limitation is the absence of explicit object grounding, causing policies to exploit spurious contextual cues rather than task-relevant visual evidence. To address this issue, we propose the Foveated Diffusion Policy (FDP), which integrates object-centric visua
Frontiers in Robotics and AI 23h ago Regulation

Imprecise beliefs: a tiny introduction

Alignment Forum yesterday

Examining the Roles of Technology Across the Health Care Journey for Individuals With Obsessive-Compulsive Disorder: Qualitative Interview Study

Background: As digital technologies become increasingly embedded in daily life, their roles in mental health care have expanded and diversified. Digital tools are being explored as interventions for obsessive-compulsive disorder (OCD) across the care continuum, including symptom recognition, access to care, treatment, and self-management. However, there is limited empirical understanding of how individuals living with OCD use digital technologies in situ to navigate their health care journeys or
JMIR (Journal of Medical Internet Research) yesterday Healthcare

Experts disagree on how to fight AI disinformation, but agree that health and politics need different solutions

When 54 international experts assessed AI-generated disinformation threats, they revealed a surprising pattern: while video deepfakes received the highest average threat ratings in the political domain (M = 6.31/7), the pattern differed in the health domain, where AI-generated text received the highest average rating (M = 5.80). The post Experts disagree on how to fight AI disinformation, but agree that health and politics need different solutions first appeared on HKS Misinformation Review .
HKS Misinformation Review yesterday MisinformationHealthcare

Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study

Background: Hypospadias is a common congenital malformation requiring surgery. Caregivers face substantial perioperative information needs, and large language models (LLMs) offer a potential health education channel, but their performance in pediatric urology and the relation between citation accuracy and clinical content safety lack systematic evaluation. Objective: This study aimed to evaluate 5 LLMs (ChatGPT-4o, Gemini-2.5-Pro, OpenEvidence, Zhipu Qingyan, and DeepSeek) for pediatric hypospad
JMIR (Journal of Medical Internet Research) yesterday HealthcareChildren & education

Social Contagion in COVID-19 Discussions Within the Belgian Reddit Community: Statistical and Modeling Study

Background: Understanding how sentiment toward COVID-19 mitigation measures evolves on social networks can help to inform infectious disease models and policymakers. Even though numerous studies have described social media interactions during the pandemic, few have modeled the underlying dynamics of sentiment contagion and polarization. Objective: This study aimed to investigate topic emergence and sentiment evolution in COVID-19 mitigation discussions on r/Belgium, focusing on (1) whether discu
JMIR (Journal of Medical Internet Research) yesterday Finance, VC & PE

Detecting Narcissistic Personality Disorder Traits on Forums: Proof-of-Concept Study

Background: Identifying traits of narcissistic personality disorder (NPD) is clinically challenging, yet early detection can significantly improve outcomes. Online forums have become a major source of self-expression, offering new opportunities to understand mental health. However, analyzing this complex language requires new tools. Objective: This study aims to determine whether a machine learning model could be trained to reliably detect language patterns associated with NPD traits in Reddit p
JMIR (Journal of Medical Internet Research) yesterday Healthcare

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended
arXiv cs.AI yesterday Jobs & economyAgents & autonomy

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analysis (GCA), team surveys, and lexical analyses of team discourse. Teams completed a high-stakes moral-dilemma decision task in a randomized controlled study: 16 teams of two students plus an AI teammate, and 17 all-human t
arXiv cs.AI yesterday Children & education

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, are already known. In reality, a partner's true capabilities are often hidden, and human collaborators may act sub-optimally on tasks with multiple valid strategies. To address these limitations, we ex
arXiv cs.AI yesterday Agents & autonomy

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a pr
arXiv cs.AI yesterday Agents & autonomy

Anatomy Contextualized Adaption of CT Foundation Models

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computatio
arXiv yesterday

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with minority-class coverage dropping to as low as 0.5% on certain datasets. To characterize and address this limitation, we conduct a
arXiv cs.AI yesterday Healthcare

DLAM: Distributional Latent Actions with Temporal Constraints

Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions ma
arXiv cs.AI yesterday Agents & autonomy

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we introduce Hybrid Ontology Matching (HOM), a new OM task that unifies equivalence and subsumption discovery, and accordingly propose a Large Language Model (LLM)-based multi-agent OM framework AgentMap that is implemented
arXiv cs.AI yesterday Agents & autonomy

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20
arXiv cs.AI yesterday Healthcare

Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment

There is growing interest in using large language models (LLMs) as low-cost proxies for resident attitudes in urban planning. Previous work shows that LLMs can predict average results of survey experiments, but less is known about whether they preserve the spatially anchored, identity-conditioned structure behind those averages, namely how support changes as a project approaches homes and how that response divides across tenure and partisan groups. We compared eight open-weight LLMs with 843 res
arXiv yesterday

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. The few existing studies on scholarly charts remain confined to visual-surface comparisons, failing to verify caption alignment, citation re
arXiv yesterday Safety & alignment

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure. Routers and retrievers can rank candidate tools by relevance, but a ranking alone does not determine how many are worth selecting. Existing approaches leave acquisition under heterogeneous costs unaddressed.
arXiv yesterday PrivacyAgents & autonomy

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender ca
arXiv yesterday Regulation

MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair

Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent benchmarks increasingly examine agent memory security, yet few trace the same malicious semantics across persistence, downstream consequences, and selective repair under diverse memory-backend comparisons. To address this ga
arXiv cs.AI yesterday PrivacyAgents & autonomy

Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise

We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite $p$-th central moment for some $p \in (1, 2]$. While static regret is well-understood, achieving universal dynamic regret in a parameter-free manner remains an open challenge. We resolve this by proposing \textbf{HT-PAder}, a parameter-free algorithm combining restarted AdaGrad experts over a geometric pool of block lengths with a pathwise m
arXiv cs.AI yesterday Environment

Value Generalisation 3: Pre-aligned AIs

Alignment Forum yesterday

Value Generalisation 2: The Missing Hole in AIs’ abilities

Alignment Forum yesterday

Value Generalisation 1: a Research and Deployment Program

Alignment Forum yesterday

Visual Credit Audit for Multimodal Spatial Reasoning

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yield
arXiv yesterday Transparency

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image assessment methods to scientific figure quality assessment, limitations emerge: classic IQA models capture perceptual quality or aesthetics but cannot judge whether a figure serves the paper's scientific argument; CLIP-ba
arXiv yesterday Safety & alignment

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work,
arXiv cs.AI yesterday Agents & autonomy

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression
arXiv cs.AI yesterday Children & education

Anticipatory Data Governance in the Age of AI: Emerging Signals in Data Access, Reuse, and Sovereignty

This paper reports findings from a structured participatory foresight study comprising two expert forecasting studios convened by The GovLab between 2025 and 2026. The studios brought together nineteen senior practitioners spanning official statistics, digital and trade policy, open science, AI governance, geospatial systems, and public-sector innovation across multiple jurisdictions. Applying a qualitative signal-scanning methodology grounded in the horizon-scanning and anticipatory-governance
arXiv yesterday Regulation

SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception

Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations to significantly accelerate on-robot learning in both egocentric and exocentric visual setups. We model a Markov Decision Process (MDP) under a symmetry tree, in which state-action pairs have admissible parallelized invari
arXiv cs.AI yesterday RegulationAgents & autonomy

Progressive Multimodal Alignment for Continual Instruction Tuning

Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributions and evolving instruction semantics cause this shared projector to drift, leading to projector-level forgetting, an issue largely overlooked by methods that focus primarily on the LLM backbone. We introduce Progressive Multimodal Align
arXiv yesterday Safety & alignment

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottleneck on consumer-grade hardware, where the computational cost of tree management severely limits inference rates. Furthermore, without deep search, these models suffer from hallucination, proposing moves with high confidence that are strateg
arXiv cs.AI yesterday Regulation

What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation

Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We present a three-class detection framework distinguishing humans, bots, and AI agents, and show that the binary-vs-agent confusion is architectural: a binary human-vs-bot detector misroutes agent sessions because its label space lacks an agent class. On o
arXiv cs.AI yesterday Jobs & economyAgents & autonomy

Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning

Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense methods show limited efficacy as they overlook the deviations among benign local updates caused by statistical heterogeneity and the stealthiness of backdoor attacks. To tackle these issues, we propose FedDAB, a two-phase method that combines local contrastive regularization with alignment checking, to defend against backdoor attacks. In the first phase, FedDA
arXiv yesterday Safety & alignmentMilitary & security

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inadequate for laboratory instruments, which must be approached from their operating side while maintaining safe clearance from surrounding equipment. We introduce BioVLN, a simulation platform for developing and evaluating v
arXiv cs.AI yesterday Agents & autonomyEnvironment

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate re
arXiv cs.AI yesterday Agents & autonomy

QR Code–Enabled Self-Access Lifestyle Education in Older People Living With HIV: Pragmatic Pilot Quasi-Experimental Pretest-Posttest Implementation Study

Background: Frailty is prevalent and dynamic in older people living with HIV and is associated with adverse outcomes. Lifestyle support is recommended but difficult to deliver at scale. Digital self-access education may help, although evidence in older, multimorbid populations is limited. Objective: This study aims to evaluate 6-month changes in frailty phenotype and related outcomes after a QR code–enabled self-access lifestyle education program on Mediterranean diet and exercise routines for p
JMIR (Journal of Medical Internet Research) yesterday Children & education

Hearsay: Vision-Language Medical Diagnoses Without an Image

When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply
arXiv cs.AI yesterday Healthcare

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents

LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to act safely. We propose Think Short, Defer Smart (TSDS), a framework that synergistically integrates a lightweight convergence probe
arXiv cs.AI yesterday Agents & autonomy

ReCo: Reweighting GRPO Against Distributional Concentration

Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability. We trace this concentration to two mechanisms in the GRPO upd
arXiv cs.AI yesterday Regulation

A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities

Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate coding agents' behavior, spanning from a total ban, mandatory disclosure, to verification gates and human sign-offs. Yet, whether coding agents read and follow those rules, and behave in open source repositories, remains unknown. To estimate real-world rule compliance of coding agents, we curate 106 issues from 49 repositories containing AI contribution rules in
arXiv cs.AI yesterday RegulationMilitary & security

FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning

Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous local architectures often induce non-aligned representation spaces, making it difficult to transfer global knowledge across silos. Existing paradigms share this knowledge as model parameters, distilled predictions, or class prototypes, yet all encode it in an absolute space that must be aligned across clients. Heterogeneous backbones break this alignment, so
arXiv yesterday Safety & alignment

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first
arXiv cs.AI yesterday Agents & autonomyEnvironment

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for learning skills across tasks. SkillRise organizes related
arXiv cs.AI yesterday Agents & autonomy

Journey Operators for Structured Multi-Axis Composition

Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume. Along one axis, order matters: "the dog bit the man" is different from "the man bit the dog." Across independent axes, however, neither composition nor movement should depend on the order of axes: in an image, composing right then down should give the same result as composing down then right, and moving right then down should describe the s
arXiv yesterday

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent. We introduce a causal audit that applies controlled mes
arXiv cs.AI yesterday Agents & autonomyTransparency

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework compri
arXiv cs.AI yesterday Healthcare

ENHANCE (Tailored Intervention for Brain Health and Cognitive Enrichment for Cognitive Health), a Coach-Supported Digital Intervention for Dementia Prevention in Underserved Older Adults: Co-Design and Usability Study

Background: Digital multidomain interventions hold promise for dementia risk reduction; however, populations at higher dementia risk, including those experiencing socioeconomic and educational disadvantage, remain underrepresented in trials, and engagement with digital interventions often declines over time. Coproduction and blended models that combine digital tools with human support may improve reach, acceptability, usability, and sustained engagement. Designing interventions that are usable a
JMIR (Journal of Medical Internet Research) yesterday Healthcare

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. In this work, we propose a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image. A lightweight Face Adapter, together with soft-tuning of the face encoder's upper blocks, aligns face-recognition features with the style space of a frozen StyleT
arXiv yesterday

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target images without extensive prompt engineering. However, existing prompt inversion methods suffer from significant limitations: (1) gradient-based methods are unstable and uninterpretable, often resulting in generated images with severe artifacts; (2) gradient-free methods yield human-readable prompts but still fail to preserve visual fidelity due to the lack of
arXiv yesterday

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for conditioning; however, implicit linguistic information from V/UV tags cannot guarantee lyric accuracy, leading to a high phoneme error rate (PER). Inspired by singing voice synthesis (SVS), we propose MPEcho, which integrates a phoneme encoder and a len
arXiv yesterday

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-cor
arXiv yesterday Agents & autonomy

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce
arXiv yesterday Safety & alignment

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconst
arXiv yesterday Agents & autonomyEnvironment

FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking

Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI models can now inject localized, high-fidelity manipulations, creating deceptive attacks that bypass standard verification. Training robust image forensic models to detect these anomalies is hindered by privacy regulations, forcing reliance on synthetic templates lacking the intricate visual patterns of real IDs. To bridge this domain gap, we intr
arXiv yesterday RegulationPrivacy

Contrastive ESA: Human Evaluation of Multiple Translations at Once

Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to
arXiv cs.HC yesterday Children & education

RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment

Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. Yet, existing deep learning approaches require dataset-specific training, large labeled corpora, and repeated adaptation to new sensor settings or activity taxonomies. Retrieval-Augmented Generation for Human Activity Recognition (RAG-HAR) addresses this by framing HAR as a training-free, retrieval-augmented task, in which statistical descriptions
arXiv cs.LG yesterday PrivacyHealthcare

WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models

Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their adoption as backbones for foundation recommendation models (FRMs). Existing approaches typically enhance recommendation with explicit Chain-of-Thought (CoT) under the Think-then-Answer paradigm. However, generating lengthy rationales introduces substantial inference overhead, while fixed CoT templates struggle to model diverse, dynamic, and context-dependent user interests. We propose WhisperRec, an ef
arXiv yesterday

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request in isolation within the current coding session, often through eliciting additional clarification. However, whether resolved session history from the same user can serve as memory for resolving recurring personalized ambiguity in a
arXiv yesterday

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterfactual reasoning, and automating simulation workflows for scientific research. In this paper, we propose Eco3S, a socio-economic system simulation framework for economic research and policy analysis that addresses these challenges through three key
arXiv yesterday RegulationAgents & autonomy

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap. The natural fix is a guard-agnostic recover-and-decode amplifier that transcribes image content and restates encoded text into its plain payload before the guard, so any off
arXiv red teaming query yesterday Safety & alignmentMilitary & security

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. Existing platforms typically deploy each workflow as an opaque GPU function, provisioning, placing, and scaling all constituent models in the workflow together. This monolithic design obscures workflow structure, inflates scaling overhead, forces users to manage low-level GPU coordination, and limits fine-grained fairness in multi-tenant
arXiv yesterday Bias & fairness

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with pr
arXiv red teaming query yesterday Safety & alignment

A Design Study on Voice-based Interaction for Immersive Network Visualization and Analysis

Visual network analysis leverages network visualization authoring techniques to facilitate sensemaking, serendipitous discovery, and hypothesis verification on network data. However, transferring the same paradigm to immersive environments is non-trivial due to insufficient UI affordance for authoring operations. Researchers have studied combining multiple modalities for interactions, but the high learning curve of such input systems limits their adoption by typical data analysts, let alone for
arXiv cs.HC yesterday Environment

Parameterized Fair Resource Allocation under Diversity Constraints

Resource allocation across multiple agent groups arises in many applications including e-commerce recommendation systems, housing assignment, and course allocation, and is commonly formulated as an optimization problem with diversity constraints to ensure group fairness. Existing approaches typically enforce these constraints as hard conditions, which overly restrict the feasible solution space and often lead to suboptimal allocations. In this paper, we propose PRA, a parameterized framework for
arXiv fairness query yesterday Bias & fairnessAgents & autonomy

Navigating the DEIverse: A comprehensive review and research agenda on diversity, equity, and inclusion in the metaverse

Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Paloma Almodóvar, Alberto Ferraris
Technology in Society yesterday Bias & fairness

Reconfiguring the global factory: The synergistic role of additive manufacturing, explainable AI, and knowledge acquisition in MNC operations

Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Femi Olan, Konstantina Spanaki, Uchitha Jayawickrama
Technology in Society yesterday Transparency

Artificial Intelligence Ethical Awareness of University Students in Ghana: A Network and Latent Profile Analyses

Publication date: Available online 27 July 2026 Source: Computers and Education: Artificial Intelligence Author(s): Bernard Yaw Sekyi Acquah, Iddrisu Salifu, Francis Arthur, Mark Inkoom, Sheriff Dagimah Suradji, Christian Inkoom, Emmanuel Quayson, Silas Afutu Quaye, Francis Obeng Gyedu, Sharon Abam Nortey
Computers and Education: Artificial Intelligence yesterday Children & education

From Simulation to Flight: Simulation-Assisted Drone Learning with Teacher-AI Co-Designed Scaffolds for Secondary Students’ STEM Knowledge and Competencies

Publication date: Available online 27 July 2026 Source: Computers and Education: Artificial Intelligence Author(s): Richard Chung Yiu Yeung, Chi Ho Yeung, Daner Sun, Therese Keane, Yuqin Yang
Computers and Education: Artificial Intelligence yesterday Children & education

How experience moderates the impact of AI suggestions on researchers' perceptions of their ideas

Publication date: October 2026 Source: Research Policy, Volume 55, Issue 8 Author(s): Matthias Tröbinger, Anil R. Doshi, Sen Chai
Research Policy yesterday Regulation

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters,
arXiv yesterday Agents & autonomy

Simulating Single Transferable Voting for the Colorado House of Representatives

arXiv:2607.25105v1 Announce Type: new Abstract: Social choice theory research demonstrates that single transferable voting (STV) results in more proportionally representative legislative bodies. We aim to understand how using multi-member districts and ranked ballots with STV would affect the representation of political parties in the Colorado House of Representatives. We investigated this objective by producing 10,000 multi-member districting plans of Colorado, generating ranked ballots for eac
arXiv cs.CY yesterday Finance, VC & PE

Passive wearable physiology tracks a state-level material-hardship gradient in resting heart rate

arXiv:2607.25301v1 Announce Type: new Abstract: Resting heart rate is an established marker of cardiovascular risk, but population-scale measurement has depended on clinical or survey instruments. We ask whether passively sensed consumer-wearable physiology recovers the socioeconomic gradient established in clinical cohorts. Using 19.1 million quality-filtered photoplethysmography readings from 18,734 opt-in users of the Welltory app, we computed cohort-adjusted mean daytime resting heart rate p
arXiv cs.CY yesterday Healthcare

Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data

arXiv:2607.25526v1 Announce Type: new Abstract: How should researchers measure the geopolitical preferences expressed by large language models (LLMs)? Existing audits commonly rely on surveys and simple tests, but international-relations research has long recognized that measuring geopolitical preferences is difficult and has developed methods for recovering them from observed choices. This paper applies a dynamic ordinal ideal-point approach from international relations, treating LLMs as respon
arXiv cs.CY yesterday Transparency

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing

arXiv:2607.25648v1 Announce Type: new Abstract: Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - under
arXiv cs.CY yesterday Regulation

Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course

arXiv:2607.24755v1 Announce Type: cross Abstract: This full research paper examines how different forms of learner-AI interaction relate to learning outcomes in object-oriented programming (OOP) courses. Generative artificial intelligence (GenAI) tools are increasingly used by students in programming education, yet evidence on their educational impact remains mixed. In particular, little is known about how students integrate GenAI tools when learning OOP, and how different patterns of use relate
arXiv cs.CY yesterday Children & education

From Idea to Classroom in Days: Using "Vibe Coding" to Create a Programming Process Visualizer from IDE Activity Logs

arXiv:2607.24757v1 Announce Type: cross Abstract: This paper reports on the rapid development and classroom deployment of a Thonny log visualizer built using AI-assisted ``vibe coding'' to make students' programming processes easily visible to teachers. We developed a web application that analyzes log files generated by Thonny (an IDE for Python) and produces interpretable views of students' programming processes. Teachers can upload a log, a ZIP archive, or a folder containing logs for a group
arXiv cs.CY yesterday Children & education

Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents

arXiv:2607.24759v1 Announce Type: cross Abstract: Research projects, educational efforts, and adjacent knowledge work accumulate findings, decisions, and reasoning that future collaborators rarely recover. The parts most useful to that work, including dead ends and walked-back claims, are routinely excluded from publications and shared code; future researchers re-attempt the same failures because no record survives. LLM coding agents are common participants but hold no persistent memory across s
arXiv cs.CY yesterday Agents & autonomy

PATHFinder Agent for Tailored Prenatal Care

arXiv:2607.24768v1 Announce Type: cross Abstract: Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through
arXiv cs.CY yesterday HealthcareAgents & autonomy

Empathy and the Human-Moment Gaps of AI Chatbots: Insights from Empathy Displacement Theory

arXiv:2607.24775v1 Announce Type: cross Abstract: Artificial intelligence (AI) chatbots are increasingly deployed in domains where empathy is essential, including healthcare, education, and customer service. However, their capacity to sustain authentic human moments remains structurally limited. This paper introduces two interlinked conceptual models to explain and address this limitation. First, the Human-Moment Gap Framework (HMGF) identifies three structural empathy deficits in AI-mediated in
arXiv cs.CY yesterday Jobs & economyHealthcare

The AI Wave and the Reinvention of Game Discovery: Oversupply, Structural Correction, and Agentic Player-Game Matching

arXiv:2607.25010v1 Announce Type: cross Abstract: AI-assisted production has sharply reduced the cost and team size required to ship a video game, producing a supply shock on open marketplaces. Recent estimates put Steam release volume at roughly sixty new titles per day, with median per-title revenue for a large share of releases falling below the platform's own submission fee [1]. This paper asks whether the resulting oversupply constitutes an emerging market crash or a structural correction,
arXiv cs.CY yesterday Agents & autonomy

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being

arXiv:2607.25057v1 Announce Type: cross Abstract: As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention. While consumer-facing generalist models can provide benefits, including improved access to information, learning, productivity, self-reflection, and companionship, they also introduce risks, such as emotional entanglement, unhealthy dependence, and the amplification of psychological vulnerabilities. Dr
arXiv cs.CY yesterday Jobs & economy

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play

arXiv:2607.25425v1 Announce Type: cross Abstract: Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper report
arXiv cs.CY yesterday Bias & fairness

Detecting CSAM Text-to-Image LoRAs From Weights

arXiv:2607.25750v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) fine-tuning has made it cheap and easy to customize open-weight image generation models for specific tasks, including the production of child sexual abuse material (CSAM). Existing moderation relies on metadata or generated outputs, but metadata can be deceptive and generating outputs may itself be unacceptable or illegal. We show that a safer signal lives in the weights. The top-left singular vectors of a LoRA's update
arXiv cs.CY yesterday Children & education

Effort Matters in Score-Based Admissions: How Retaking and Aggregation Shape Test Scores

arXiv:2607.25974v1 Announce Type: cross Abstract: Observed standardized test scores are the result of an endogenous process: students strategically allocate effort across multiple retake attempts to improve their outcomes. Because students differ in their ability to make these investments, the interaction between applicant strategy and institutional scoring rules---such as the widely used Single-Sitting and Superscoring policies---can disparately distort observed scores. We develop a strategic f
arXiv cs.CY yesterday Children & educationFinance, VC & PE

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

arXiv:2507.11548v3 Announce Type: replace Abstract: The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias relative to human judgment. However, this framing leaves a prior question unresolved: whether these systems are capable of performing the evaluative task at all. This study presents a two-part audit of eight widely used AI platforms used for resume screening. Drawing on the concept of the Illusion of Neutra
arXiv cs.CY yesterday Bias & fairnessTransparency

Three Lessons from Citizen-Centric Participatory AI Design

arXiv:2602.08554v2 Announce Type: replace Abstract: This workshop paper examines challenges in designing agentic AI systems from a citizen-centric perspective. Drawing on three participatory workshops conducted in 2025 with members of the general public and cross-sector stakeholders, we explore how societal values and expectations shape visions of future AI agents. Using constructive design research methods, participants engaged in storytelling and lo-fi prototyping to reflect on potential commu
arXiv cs.CY yesterday Agents & autonomy

LLM-generated personalized nudges for improving pro-environmental behavior: Field evidence from resource conservation

arXiv:2604.03881v2 Announce Type: replace Abstract: Encouraging pro-environmental behavior remains a major challenge for sustainable cities. Conventional feedback nudges can show individuals how their current behavior compares with environmental goals but often provide limited guidance on what to do differently in daily life. This study examines whether supplementing weekly feedback on participants' behavior with LLM-generated personalized action suggestions improves pro-environmental behavior,
arXiv cs.CY yesterday Environment

Less Deliberate in Teams: Student LLM Use Across Individual and Collaborative Work

arXiv:2606.30860v2 Announce Type: replace Abstract: As large language models (LLMs) become common in computing courses, we need to understand how the social setting shapes how students use them. This paper reports findings from a semester-long study of 96 undergraduate students who completed six assignments, alternating between individual homework and team project milestones. We tracked LLM usage, prompting habits, and how students verified AI-generated output across all six assignments. LLM usa
arXiv cs.CY yesterday Children & education

Generative AI Availability, Grades, and Student Satisfaction at a Large University

arXiv:2607.21534v2 Announce Type: replace Abstract: The spread of generative AI (GenAI) in higher education has raised concerns that students offload cognitive effort to AI, earning high grades without learning. If this "GenAI substitution hypothesis" is true, grades should rise disproportionately in GenAI-susceptible courses--those relying more on assessments like take-home problem sets and essays rather than in-class exams. Substitution could also affect student satisfaction, measured here as
arXiv cs.CY yesterday Children & education

"We'll have to see how it works": An interview study to understand collaborative practices in interdisciplinary artificial intelligence and healthcare research

arXiv:2311.18424v3 Announce Type: replace-cross Abstract: Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients and other stakeholders together. By understanding AI as 'sociotechnical' where the social and the technical nature of the work and the models are inseparable, we explore the AI development workflow and how stakeholders navigate the challenges and tensions of sharing and generating knowledge across dis
arXiv cs.CY yesterday Healthcare

When Algorithms Meet Artists: Semantic Compression and Stake-holder Marginalisation in Public AI-Art Discourse (2013-2025)

arXiv:2508.03037v5 Announce Type: replace-cross Abstract: Artists occupy a paradoxical position in generative AI. Their own work trains models that now compete with them, replicate their styles, and reshape the creative economy they inhabit. Yet whether artist concerns achieve proportional representation in the public discourse that shapes AI governance remains an open empirical question. We mapped the semantic landscape of public AI-art discourse from 2013 to 2025, drawing on 1,736 text chunks
arXiv cs.CY yesterday RegulationJobs & economy

News (283)

OpenAIs KI-Agent knackte nicht nur Hugging Face – was genau passierte

OpenAI-Modelle attackierten vergangene Woche scheinbar weitere Software. Die Firma erklärt nun ausführlicher, was genau passiert ist.
Heise Online (DE) 18h ago Agents & autonomy

China’s renewable energy generation surpasses 40% of total power output for first time in H1 2026

According to CCTV News, the National Energy Administration (NEA) held a press conference on Wednesday, announcing that China’s renewable energy sector saw rapid growth in the first half of 2026, with renewable power generation accounting for more than 40% of total electricity generation for the first time. During the first half of the year, China’s […]
TechNode (CN) 19h ago Environment

Finance firms set to pour more investment into AI amid ‘data divide’ fears

A majority of surveyed global asset management firms plan to raise their artificial intelligence budgets by at least 50 per cent within the next year as the technology transforms the finance industry, a recent study showed. The adoption of AI is also going to have a big impact on labour-intensive operations: 62 per cent of polled fund managers expect transformative change in data generation and summarisation, according to a study released Tuesday by US fintech firm Clearwater...
SCMP Tech (HK/CN) 19h ago Jobs & economyFinance, VC & PE

AI investment concentration risk is not just in equities

Bond markets are increasingly dominated by a bet on the same thesis as other asset classes
Financial Times Technology (headlines) 19h ago Finance, VC & PE

The UK has worse mobile coverage than Romania. Why?

Operators blame planning rules and low prices for a lack of investment
Financial Times Technology (headlines) 19h ago Finance, VC & PE

CuspAI's Max Welling: ‘We are building molecules to remove forever chemicals from water’

The co-founder of the UK-based science start-up explains how AI can help create new materials to address some of the world’s most complex challenges
Financial Times Technology (headlines) 19h ago Environment

Home water harvesting is being supercharged by Nobel Prize-winning tech

Start-up Ahbstra applies MOF technology to draw moisture from the atmosphere at volume, using minimal energy
Financial Times Technology (headlines) 19h ago Environment

Final Call: Age Verification and Restricting Social Media for Children, Delhi, 31 July #NAMA

Join experts at MediaNama event to discuss age verification, social media restrictions for children, platform design, privacy, online safety, and practical regulatory approaches. The post Final Call: Age Verification and Restricting Social Media for Children, Delhi, 31 July #NAMA appeared first on MEDIANAMA .
MediaNama (IN) 19h ago RegulationPrivacy

Leopold Aschenbrenner’s Situational Awareness seeks to raise capital after AI rout

Hedge fund has held talks with existing investors and lenders in recent days
Financial Times Technology (headlines) 19h ago Finance, VC & PE

Joyce Carol Oates Defends ‘The Odyssey’ and Slams Translator for Scathing Review: ‘Speaks in the Crude Language of MAGA Folks’

Author Joyce Carol Oates came to the defense of Christopher Nolan’s “The Odyssey” after translator Emily Wilson wrote a viral review attacking the film. “Rather than disagreeing with interpretations of Homer in a collegial manner, this person, who has benefited enormously from Nolan’s film, speaks in the crude language of MAGA folks attacking someone with […]
Variety (AI) 20h ago Military & security

China threatens retaliation against U.S. humanoid robot ban, says it 'severely damages' relations

China's Commerce Ministry said Thursday the U.S. Federal Communications Commission has repeatedly ignored Beijing's restrained stance.
CNBC Technology 20h ago Agents & autonomy

Zuckerberg lays out Meta's AI capacity dilemma: What to sell vs. what to keep

For investors anxious to hear more about Meta's plans to make money from its big AI spending, Mark Zuckerberg said there's a trade-off.
CNBC Technology 22h ago Finance, VC & PE

Justin Bieber’s Sneakers Worn During World Cup Final Performance Sold at Auction for Almost $30K

Christie's auction house handled the One Goal auction, which benefits the FIFA Global Citizen Education Fund, a $100 million initiative to expand access to quality education and football for children worldwide.
Hollywood Reporter (AI/entertainment) 22h ago Children & education

SpaceX faces House Energy Committee demand to tour its AI data centers in Memphis

The ranking member of the House Committee on Energy is demanding SpaceX records and a tour of its data centers and power plants in and around Memphis.
CNBC Technology 23h ago Environment

Elon Musk’s xAI sues Minnesota over law banning ‘nudification’ technology

First-in-nation law sets up test on states’ power to regulate use of AI as it tries to outlaw fake nude images of real people Elon Musk’s company xAI has sued Minnesota over the state’s first-in-the-nation law banning “nudification” technology on websites and apps, potentially providing a test for how far states can go in constitutionally regulating the use of artificial intelligence. Musk’s company sued on Monday in federal court, days before the law is set to take effect on Saturday and make M
The Guardian 23h ago Regulation

Freehand raises $75m for AI agents that manage supply chain spend

Freehand, a startup building AI agents that manage supply-chain spend for giants including Meta and Pfizer, has raised $75 million in funding.
Finextra AI 23h ago Agents & autonomy

Meta shares tumble as Zuckerberg tries to sell his vision for AI ‘agents’

Social media chief defends strategy based on personalised bots as costs rise and revenue projections disappoint
Financial Times Technology (headlines) 23h ago Agents & autonomy

White House’s new high-risk life sciences policy calls for monitoring AI dangers

“The White House Office of Science and Technology Policy (OSTP) will convene an interagency group to monitor advancements at the intersection of biological sciences and artificial intelligence, including in silico life sciences research,” the guidance says.
NextGov/FCW 23h ago RegulationBiotech

Production Assistants on CBS’ ‘Cupertino’ Vote to Unionize

The New Jersey-based workers unanimously declared their support for joining LiUNA Local 724 in the labor group’s first victory on a CBS Studios series.
Hollywood Reporter (AI/entertainment) 23h ago Jobs & economy

Digital avatar of Jair Bolsonaro tests Brazil’s AI election rules

Lawyers for former president insist he did not authorise the projection following legal challenge from leftwing groups
Financial Times Technology (headlines) 23h ago Misinformation

OpenAI CFO Sarah Friar tells employees that annualized revenue in July topped all of Q2

OpenAI is trying to reassure employees that the business is healthy as competition emerges from Anthropic as well as a host of open-source players.
CNBC Technology yesterday Healthcare

K-12 cyber information exchange, incident catalog would be directed by CISA under new legislative proposal

The Enhancing K-12 Cybersecurity Act failed to pass in 2023. Sens. Mark Warner and Marsha Blackburn said their bill is urgently needed in 2026.
StateScoop yesterday RegulationMilitary & security

Mark Zuckerberg predicts that billions of people will have personal AI agents in five years

As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the price.
TechCrunch yesterday Agents & autonomyFinance, VC & PE

Microsoft logs $3.2B from Anthropic investment, but OpenAI was a mixed bag

When Microsoft reported killer fourth-quarter earnings for its fiscal 2026 year (which ended June 30), it tucked in an interesting little tidbit about how its investments in the two biggest, and competing, AI labs are doing.
TechCrunch yesterday Finance, VC & PE

Howard Revisits Disenrollment Decision

Howard Revisits Disenrollment Decision Joshua.Bay Wed, 07/29/2026 - 06:27 PM In a conversation with  Inside Higher Ed , Interim President Wayne Frederick said Howard reinstated more than 200 students and outlined changes to improve communication. Byline(s) Joshua Bay
Inside Higher Ed Tech yesterday Children & education

Zuckerberg says Meta’s enterprise AI opportunity extends beyond agents

On the company’s second-quarter earnings call Wednesday, CEO Mark Zuckerberg said Meta sees a “large enterprise opportunity” spanning AI agents, APIs, compute, and internal software.
TechCrunch yesterday Agents & autonomy

Microsoft confirms Copilot ‘super app’ coming this year

Microsoft is working on an AI "super app" that combines Copilot's chat, coding, and agentic capabilities. During an earnings call on Wednesday, Microsoft CEO Satya Nadella said the app will span "both consumer and commercial experiences" when it launches this year. "Copilot is evolving rapidly from chat to Cowork to Autopilots," Nadella said. "This quarter, […]
The Verge yesterday Agents & autonomy

Microsoft’s cloud just hit a new milestone—Azure crosses $100 billion in annual revenue

Microsoft’s fiscal Q4 revenue of $90 billion topped Wall Street targets, while Copilot reached 30 million paid seats, and an investment in Anthropic delivered a $3.2 billion windfall.
Fortune AI yesterday Finance, VC & PE

Mark Zuckerberg is planning a big push into personal AI agents

Meta is all-in on AI, and sometime soon, the company is going to make a big push into personal AI agents that can do things on your behalf. On Wednesday's Q2 2026 earnings call, CEO Mark Zuckerberg previewed a high-level vision of how the company is thinking about personal agents and what it will do […]
The Verge yesterday Agents & autonomy

John Thune Isn’t Bowing to Trump’s Pressure: ‘Show Me How This Ends’

A standoff over the restrictive voting bill Trump is pushing has left nerves raw at the Capitol, with even the typically staid Thune showing signs of frustration.
Time Tech yesterday Regulation

In an era defined by AI, can work still be measured by the clock?

Hwaseong, South Korea, past midnight. An engineer watches a defect pattern surface on her screen – the kind that, caught now, saves a week of wafers. She is locked in, fully absorbed, the way engineers get when a breakthrough is close. But she has to go home. The law says so. Yet something is disrupting the clock she races against. Artificial intelligence is quietly rewriting what her hours are worth. On June 29, South Korean President Lee Jae Myung stood alongside the leaders of Samsung...
SCMP Tech (HK/CN) yesterday Regulation

Discover what’s next for AI, from the SaaS reckoning to the agent security gap, at TechCrunch Disrupt 2026

At TechCrunch Disrupt 2026, the AI Stage is back to dig into the single hottest topic in the community for the past few years, presented by Google for Startups.
TechCrunch yesterday Agents & autonomy

Thinking Machines co-founder Lilian Weng left the company citing health reasons, then joined OpenAI

Weng previously served as the VP of AI Safety Research at OpenAI.
TechCrunch yesterday Safety & alignmentHealthcare

xAI’s last-minute scramble to stop Minnesota’s anti-nudification app law

xAI is suing Minnesota Attorney General Keith Ellison over a law passed back in May that broadly targets "nudification" apps, claiming that the statute's punitive provisions leave the company with "no practical choice but to restrict Grok Imagine's image-editing features in various ways." The law, the company argues, violates the First Amendment. Back in January, […]
The Verge yesterday Regulation

In Diary, Fauci Bragged of Role in School Closures. At Hearing, He Took the 5th

In a three-hour hearing Wednesday, Dr. Anthony Fauci, the influential scientist who led the national response to COVID, refused to respond to questions about his role in school closures and the mask and social distancing guidelines that experts say have drastically set back students’ learning. Subpoenaed by Sen. Rand Paul of Kentucky, the Republican chair […]
The 74 (education AI) yesterday Children & education

Musk's X says Australia social media ban crackdown undermines international law

Takes aim at eSafety.
iTnews (AU) yesterday Regulation

Senators Push Federal Regulator to Prove School Bus Safety Data Is Accurate

The post Senators Push Federal Regulator to Prove School Bus Safety Data Is Accurate appeared first on ProPublica .
ProPublica (Machine Bias) yesterday RegulationChildren & education

A.I. Companies Are Recruiting Electricians and Carpenters by the Thousands

The future of artificial intelligence depends on finding more skilled humans for some very physical jobs.
The New York Times yesterday Jobs & economy

Homeland Security plans to add more AI to FOIA processing

The agency wants automation tools to improve efficiency amid an increasing volume of requests, but advocates of transparency and accountability have concerns. The post Homeland Security plans to add more AI to FOIA processing appeared first on FedScoop .
FedScoop yesterday Jobs & economyTransparency

Who wins and who loses after US bans foreign robots?

Government ban on foreign-made robots may hinder instead of help US robotics.
Ars Technica yesterday Agents & autonomy

The Silicon Valley Health Trend Making Doctors Nervous

Is more data about your body actually good for you?
The New York Times yesterday Healthcare

OpenAI's Rogue Model Claims More Victims Beyond Hugging Face

OpenAI revealed rogue AI models compromised more services than initially disclosed, including a Modal customer environment and others.
Dark Reading (AI security) yesterday Environment

The US government's robot ban also includes vacuums

Robots apparently don't need to be bipedal to be caught up in the FCC's new foreign-made robot ban.
Engadget AI yesterday Agents & autonomy

Red Agents vs. Blue Agents: How to Make AI Better At Defense

The agentic AI playing field was heavily tilted toward offense, so researchers began using red team agents to help teach their blue counterparts.
Dark Reading (AI security) yesterday Safety & alignmentMilitary & security

Trump Administration Is Repurposing Federal Land for A.I. Data Centers

In the latest example, the Energy Department will convert a shuttered Cold War-era uranium enrichment facility into a data center campus and gas plants.
The New York Times yesterday Environment

Waymos are starting to run on freeways again

Waymo paused highway operations in May after several robotaxis drove into sections that were closed for construction.
Engadget AI yesterday Agents & autonomy

New Senate bill would update state education, workforce data systems for AI era

New bipartisan legislation would employ artificial intelligence to improve the longitudinal data systems built by state governments to improve educational and workforce outcomes.
StateScoop yesterday RegulationJobs & economy

The True Story Behind Netflix's The Idaho Murders: College Nightmare

As Bryan Kohberger challenges his conviction for murdering four Idaho college students, a new Netflix documentary examines what we know about the case
Time Tech yesterday Children & education

AI hackers are getting faster. The government may not be ready

5 current and former government technology officials assess whether federal cyber defenses can keep up.
Fast Company yesterday Military & security

AI hackers are getting faster. The government may not be ready

Agentic AI systems threaten cybersecurity as we know it. A new generation of cyber-capable models, including Anthropic’s Mythos and OpenAI’s GPT-5.6, can find and exploit vulnerabilities in computer systems far faster than human hackers. The technology is so powerful that it has spooked the U.S. government, which has moved to limit, or completely pause, the public release of these models. How well prepared is the Trump administration to secure the government’s computing resources? The rise of ag
Fast Company Tech yesterday Military & securityAgents & autonomy

Anthropic backs urgent call for the most powerful AI labs to hit the brakes

Less than a week after OpenAI disclosed that two experimental AI models escaped their testing environment during a cybersecurity exercise The post Anthropic backs urgent call for the most powerful AI labs to hit the brakes appeared first on The New Stack .
The New Stack AI yesterday Environment

“The beast needs a cage”: Why PortSwigger’s agentic pentesting is kept safe behind bars

As agentic services diversify across the entire enterprise technology stack, the rise of agentic coding tools is being challenged by The post “The beast needs a cage”: Why PortSwigger’s agentic pentesting is kept safe behind bars appeared first on The New Stack .
The New Stack AI yesterday Agents & autonomy

It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

I watched a new tool try to get around the model safeguards of four major frontier companies. You might be surprised by how they performed.
Wired yesterday Safety & alignment

You’re Not Imagining It: Your Kid’s School Supplies Cost More This Year

Pencils are more expensive. So are notebooks, scissors, glue sticks — just about every back-to-school item that families will load into their literal or virtual carts this summer has gone up in price at a time when American families already feel crushed by rising costs. According to one analysis of the 21 most common school […]
The 74 (education AI) yesterday Children & education

Kenya empowers investigators to seize crypto wallets tied to financial crime

Kenya's new crypto rules now empower authorities to seize and convert crypto assets into fiat asset to preserve value during investigations.
TechCabal (Africa) yesterday Finance, VC & PE

Sam Altman previews new AI model on Capitol Hill after cyber breach

The OpenAI CEO discussed a forthcoming artificial intelligence model with federal lawmakers as the cybersecurity debate over the advanced technology intensifies.
Politico Technology (US) yesterday Military & security

PwC has allegedly published AI-generated reports containing false or fabricated sources

Following KPMG, Deloitte, and Ernst & Young, GPTZero has now found fabricated sources and false claims in four PwC Middle East reports. One governance report scored 84 percent AI-generated and promoted a PwC product with unverified customer references. All Big Four firms are now affected by AI hallucinations. The article PwC has allegedly published AI-generated reports containing false or fabricated sources appeared first on The Decoder .
The Decoder yesterday Regulation

Who's Liable When AI Agents Escape? Hugging Face Breach Raises Hard Questions

Dark Reading walks through the many twists and turns in the bizarre story of how OpenAI's agent AI system broke out of its sandbox and decided to target Hugging Face, and what CISOs should be aware of.
Dark Reading (AI security) yesterday Agents & autonomy

Hugging Face Hack Lessons for Cyber Defenders

Dark Reading Confidential Episode 20: Expert Rich Mogull reflects on lessons cyber teams should pull from the OpenAI agent's attack on Hugging Face.
Dark Reading (AI security) yesterday Military & securityAgents & autonomy

How Mexico Locked In Its Militarized Police State

In 2021, Andrew Ivey wrote, “‘I Have Other Data’: The Guardia Nacional and the Entrenchment of Mexico’s Militarization’,” where he warned Mexico was headed toward a militarization trap. Five years later, amidst increased pressure from the Trump administration and the election of a new Mexican president, we asked Andrew to revisit his arguments.Image: Sitio Oficial de Andrés Manuel López ObradorIn your 2021 article, you warned that Mexico was headed toward an entrenched militarization trap. In 20
War on the Rocks yesterday Misinformation

Kenya’s new crypto rules could force exchanges to delist foreign stablecoins

The new rules give the central bank direct oversight to cut off local access to foreign stablecoins without having to regulate the offshore issuers themselves.
TechCabal (Africa) yesterday Regulation

OpenAI says rogue agent behind Hugging Face hack broke into additional services

The four additional targeted organizations weren’t named. OpenAI said they were not affected as severely as Hugging Face.
The Record (Recorded Future News) yesterday Agents & autonomy

Mova presenta su primer localizador GPS para mascotas: habla con tu perro a distancia y rastréalo en tiempo real

Los dispositivos de localización de objetos tipo AirTag han ganado mucha popularidad en los últimos años para seguirle la pista a todo tipo de pertenencias. Sin embargo, cuando se trata de animales, los sistemas pasivos por Bluetooth se quedan cortos en zonas de campo o cuando el animal se desplaza rápido. Mova, una marca reconocida en el sector del hogar conectado por sus robots aspiradores , ha dado el salto al cuidado de animales con el SureTrack Pro , un localizador GPS diseñado específicame
Xataka (ES) yesterday Agents & autonomy

Llevamos décadas dando por hecho que la expansión del universo se acelera por la energía oscura. Un estudio dice que todo fue un espejismo estadístico

Los astrónomos llevan décadas intentando desentrañar los misterios de la materia y la energía oscura . Sin embargo, ahora un equipo de científicos de la Universidad de Oxford ha lanzado la más controvertida de las soluciones : no hay nada que buscar, porque la energía oscura no existe. Y se han quedado tan panchos. Lógicamente, las reacciones no se han hecho esperar. Algunos investigadores creen que es una solución interesante en la que vale la pena seguir indagando. Otros, en cambio, consideran
Xataka (ES) yesterday Finance, VC & PE

The Lone Republican Who Voted Against Advancing Bill to Put Tougher Sanctions on Russia

The lawmaker referred to the legislation as the "latest counterproductive attempt to hold Russia accountable for its war against Ukraine."
Time Tech yesterday RegulationTransparency

What X's Corrected Action Under the DSA Fixes on Data Access, and What's Left Open

Tech Policy Press yesterday

Advanced AI Is Ultrahazardous. Let’s Treat It That Way

Tech Policy Press yesterday

MoonPay opens PayBox for AI agent transactions

MoonPay, the global financial technology company powering the movement of value across fiat and digital assets, has launched PayBox, the first payment vault built for AI that lets a person's AI agent transact on the open internet without leaving the conversation.
Finextra AI yesterday Agents & autonomy

Navy using LETHALITY Consortium to advance network consolidation initiative

The initiative is intended to modernize and consolidate IT and data architectures at the Naval Surface Warfare Center Corona Division. The post Navy using LETHALITY Consortium to advance network consolidation initiative appeared first on DefenseScoop .
DefenseScoop yesterday Military & security

New Staffing Models Put Strong Teachers Out Front

In Carlsbad, N.M., a tiny district in the southeastern corner of the state, nine of its 12 schools have turned teacher assignments upside-down, inviting their best teachers to coach a team of colleagues, model lessons and take direct responsibility for the achievement of every student the team serves. In exchange, each teacher-leader earns substantially more […]
The 74 (education AI) yesterday Children & education

Why trust has become Nigeria’s next digital payments challenge

Nigeria no longer has a payments problem. It has a trust problem. That is the central argument of a new report by Bridgforte, a policy research institute, published in partnership with the United Nations Development Programme (UNDP) Innovation Hub in Lagos, which argues that confidence in Nigeria's financial system is now determined less by how quickly money moves than by how reliably the system responds when something goes wrong. The country processed more than ₦1.2 quad
TechCabal (Africa) yesterday Regulation

Rogue OpenAI agent compromised second tech firm's customer

An OpenAI agent compromised a customer of another technology company, the New York-based firm Modal Labs announced Wednesday. In a technical timeline posted Tuesday, the tech startup Hugging Face explained how an OpenAI agent escaped the AI firm's isolated testing sandbox and accessed another testing environment "hosted by a user of a third-party infrastructure provider."...
The Hill Technology yesterday Agents & autonomyEnvironment

Students: Is A.I. Changing Your Life? Tell Us.

We’re looking for college and high school students to tell us about A.I.
The New York Times yesterday Children & education

Mail spent £35m fighting Harry’s ‘campaign for Leveson 2’ privacy claim

Claim was launched with 'blaze of publicity' and 'monstrous all-out attack' on Mail. The post Mail spent £35m fighting Harry’s ‘campaign for Leveson 2’ privacy claim appeared first on Press Gazette .
Press Gazette (AI/journalism) yesterday Privacy

Jim Cramer says Seagate is the ‘key to this market’ after blowout earnings

Jim Cramer said investors should look to Seagate Technology as a barometer of whether the AI infrastructure trade can regain momentum on Wednesday.
CNBC Technology yesterday Finance, VC & PE

Nigeria tests investor appetite with sale of former telecom monopoly ntel

The process remains in its early stages. No investors have emerged yet, AMCON spokesperson Jude Nwauzor said.
TechCabal (Africa) yesterday Finance, VC & PE

Les États-Unis interdisent les robots humanoïdes chinois, au nom de la sécurité nationale

La FCC a ajouté les robots humanoïdes et les onduleurs électriques fabriqués à l'étranger à sa liste noire des équipements jugés dangereux pour la sécurité nationale américaine. Une décision qui vise sans le dire la Chine, leader mondial de la robotique avancée.
Numerama (FR) yesterday Agents & autonomy

“Stateful systems are incredibly hard to build”: How Perplexity thinks about AI agent sandboxes

Sandboxes for AI agents may feel like a solved problem. After all, projects like Firecracker, the open-source microVM technology AWS built The post “Stateful systems are incredibly hard to build”: How Perplexity thinks about AI agent sandboxes appeared first on The New Stack .
The New Stack AI yesterday Agents & autonomy

Toheeb Popoola found Photoshop by accident. It changed his life.

As his family's financial situation worsened, earning an income became just as important as pursuing a university education.
TechCabal (Africa) yesterday Children & education

The Fed meeting, Ford earnings, Trump's investment portfolio and more in Morning Squawk

Here are five key things investors need to know to start the trading day.
CNBC Technology yesterday Finance, VC & PE

Tencent-owned Lightspeed LA is laying off staff

The Last Sentinel developer attributes the job cuts to a shift in creative and development direction for the project.
Game Developer (AI) yesterday Jobs & economy

FTSE 100 hits record high despite AI sell-off

Strong corporate results buoy market as investors move money away from tech and semiconductor stocks London’s FTSE 100 stock index has touched a fresh high, driven by strong corporate results as investors moved money away from tech and semiconductor stocks amid the global tech stock sell-off . The UK’s blue chip index rose as high as 10,951 points on Wednesday morning before falling back slightly, its best level since 27 February, the day before the US and Israel began attacks on Iran and sparke
The Guardian yesterday Finance, VC & PE

Open weights vs. closed: An AI civil war's afoot, and the stakes are existential

It started as China vs. the US, but it's become a face-off between two fundamentally different ways of building LLMs. And the safety of everything is on the line.
ZDNet AI yesterday Safety & alignment

Atlassian tightens tracking of staff AI use as other technology firms encourage ‘tokenmaxxing’

Some companies reportedly using leaderboards for employees who used the most AI in their work Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast Software firm Atlassian has sought to tighten tracking of its staff’s AI spending by introducing “wallets” with monthly caps of up $2,000 for each employee, amid an explosion in costs at other tech companies. The move by Atlassian, which recently cited AI as part of the reason behind cutti
The Guardian yesterday Privacy

Compliance is becoming infrastructure, not overhead

For years, financial institutions have treated compliance as a cost of doing business. A regulation ...
Finextra AI yesterday Regulation

FCC Bans Foreign-Made Humanoid Robots, Targeting China Over National Security

China has criticized the move, accusing the U.S. of protectionism.
Broadband Breakfast yesterday Military & securityAgents & autonomy

Encore AI raises $30M to build AI agents that learn from customer calls

The startup analyzes calls, messages, and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
TechCrunch yesterday Agents & autonomy

Patch-Resistant 'RufRoot' Flaw Can Unleash Malicious AI Agent Swarms

The vulnerability in the AI hosting platform Ruflo allows an unauthenticated attacker to take over the system and corrupt memory, so bad behavior can persist after patching.
Dark Reading (AI security) yesterday Agents & autonomy

Opinion: Threat of Deportation, Program Cuts Disrupt Student Learning

Once upon a time, a knight in shining armor rode his gallant steed into a village where he was met with cheers and celebration. I like to believe that these heroes still exist — I know one named Fluviano. He’s a student in my English language development class. He does not ride atop a gallant […]
The 74 (education AI) yesterday Children & education

Trump administration bans foreign-made humanoid robots in move targeting China

The Federal Communications Commission (FCC) has banned imports of foreign-produced humanoid robotic devices and power inverters, citing “national security” as its primary concern. The devices were added to the FCC’s covered list, meaning the products posed an “unacceptable risk to the national security of the United States or the safety and security of United States...
The Hill Technology yesterday Military & securityAgents & autonomy

The latest Orange Rag Product Table – link here

If 2024 was about GenAI experimentation and 2025 was about finding use cases, 2026 is the year of operationalisation, governance and control. Take a look at our product launch table […] The post The latest Orange Rag Product Table – link here appeared first on Legal IT Insider .
Legal IT Insider yesterday Regulation

Minister apologizes as Korean leveraged ETF investors nurse heavy losses amid chip stock rout

Korean retail investors have racked up heavy losses from leveraged bets on stocks, following rule changes earlier this year.
CNBC Technology yesterday Finance, VC & PE

Siobahn Day Grady Wants Everyone to Be AI Literate

Artificial intelligence is reshaping the skills employers expect from new graduates. In response, universities are scrambling to launch new courses, research centers, and industry partnerships that prepare students for today’s workforce. But building a cutting-edge AI curriculum demands funding and access to industry networks, resources that remain unevenly distributed across higher education. At North Carolina Central University, Siobahn Day Grady is trying to change that equation. In January 2
IEEE Spectrum yesterday Jobs & economyChildren & education

How a medical database developed at MIT evolved into a global standard of data-sharing

The visionary PhysioNet platform launched 25 years ago, based on a system developed at MIT in the 1970s. It has become one of the most comprehensive biomedical and clinical data repositories in existence.
MIT News yesterday Healthcare

Exclusive: Rent the Runway cofounder Jenn Hyman is taking over IPO-bound Babylist as it nears $1 billion in revenue

Babylist founder Natalie Gordon will become executive chair, in a rare founder-to-founder CEO transition.
Fortune AI yesterday Finance, VC & PE

Turning 10x developers into 10x value

When I sit down with leaders and ask why they’re investing in AI, the answer almost always comes back to The post Turning 10x developers into 10x value appeared first on The New Stack .
The New Stack AI yesterday Finance, VC & PE

Brookfield to turn former nuclear weapons site into AI campus

Canadian investment group to partner utility NextEra on data centre in Kentucky
Financial Times Technology (headlines) yesterday Military & securityFinance, VC & PE

Künstliche Intelligenz: Hackerangriff von OpenAI umfangreicher als bislang bekannt

Der Angriff einer KI von OpenAI hat größere Ausmaße als bislang vermutet. Der KI-Agent attackierte vier weitere Onlinedienste und griff gezielt fremde Zugangsdaten ab.
Zeit Digital (DE) yesterday Agents & autonomy

SpaceXAI veut faire annuler la loi du Minnesota contre la « nudification »

SpaceXAI part en croisade contre une loi du Minnesota interdisant les technologies de « nudification », qui permettent de générer des deepfakes dénudés de personnes réelles. L’entreprise est la première concernée, puisque son chatbot Grok est massivement utilisé pour créer ces images. En début d’année, Grok était sous le feu des critiques : le chatbot de xAI […]
Next (FR, ex-INpact) yesterday Misinformation

‘Explosive demand’ for cooling continues as Europe swelters, China’s Midea says

Chinese home appliance maker Midea reported that European demand for its air conditioners continues to soar amid recurrent heatwaves, with two of its factories receiving hundreds of thousands of new orders in a month. The factories, in the eastern city of Wuhu and the southern port city Guangzhou, received new European orders for 200,000 units within a month amid “explosive demand”, the company said in a statement on Wednesday on the Shenzhen Stock Exchange’s investor Q&A platform. The company..
SCMP Tech (HK/CN) yesterday Finance, VC & PE

STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated

In this edition of AI Prognosis: A conversation about benchmarking leading clinical chatbots, investor view on AI in biopharma, and more.
STAT News (health AI, headlines) yesterday HealthcareFinance, VC & PE

Consent, Not Nudity, Should Define Image-Based Abuse Laws

Tech Policy Press yesterday

I tested a premium Linux laptop that's as light as it is powerful - here's why I love it

The Ar Gen 1 is an impressive device, suitable for anyone from students to developers to business owners.
ZDNet AI yesterday Children & education

Smarter tech investment starts with trusted data

Executives from BCX, Telkom Business and Openserve hosted a channel event to champion trusted data as the foundation of an intelligent company.
ITWeb (ZA) yesterday Finance, VC & PE

Serena Williams wants to find the next trillion-dollar company

Serena Williams doesn’t just defy convention on the tennis court. She’s also applying her competitive mindset to entrepreneurship and venture investing, recently rebranding her firm from Serena Ventures to Starfire Ventures. Williams discusses the investment strategy that has led to 16 unicorns now in her portfolio, what she looks for in founders, and how partnerships with companies from Nike to Ro fit into her broader vision.  This is an abridged transcript of an interview from Rapid Respo
Fast Company Tech yesterday Finance, VC & PE

Modus’s operandi: To give AI agents just the right amount of context

As more companies plug AI agents into the deepest depths of their internal data banks, how can they be sure The post Modus’s operandi: To give AI agents just the right amount of context appeared first on The New Stack .
The New Stack AI yesterday Agents & autonomy

Shipping code without human verification

Agents are writing code faster than humans can review it. The answer is not “review faster”; that would be like The post Shipping code without human verification appeared first on The New Stack .
The New Stack AI yesterday Agents & autonomy

heise+ | Flash-Talk: Entspannung bei Flash-Versorgung, Fälschungen, Rechenzentren im Meer

Samsung und SK Hynix investieren zusammen über eine Billion US-Dollar in neue Fabs. Silicon Motion startet die Entwicklung von SSD-Controllern für PCIe 7.0.
Heise Online (DE) yesterday Finance, VC & PE

KI-Update kompakt: KI-Allianz, Claude Cowork, MAI-Cyber, Expenditure Horizons

Das "KI-Update" liefert drei mal pro Woche eine Zusammenfassung der wichtigsten KI-Entwicklungen.
Heise Online (DE) yesterday Military & security

Pay-TV Operators Accuse FCC of Flip-Flopping on Ways to Resolve Retransmission Consent Policy Disputes

American Television Alliance says FCC is issuing conflicting guidance on raising retransmission consent cost issues in connection with TV station mergers and TV station ownership rulemakings
Broadband Breakfast yesterday Regulation

Rogue OpenAI agent that hacked startup tried to attack other firms

ChatGPT developer says activity by autonomous tool was not at severity or scale of what occurred at Hugging Face OpenAI has revealed that a cyber-attack carried out by a rogue AI agent had more than one victim. The ChatGPT developer said the agent – an autonomous tool able to carry out sequences of commands without human help – had located and used logins to access four other unnamed “publicly-available services” in addition to the US startup Hugging Face. Continue reading...
The Guardian yesterday Military & securityAgents & autonomy

US foreign robot ban to benefit Tesla, other American developers, analysts say

China’s robotics firms have run into a fresh obstacle in their global expansion, after the United States banned imports of new foreign-made models, which analysts say could benefit Tesla and other American developers. The Federal Communications Commission (FCC) added foreign-produced “advanced robotic devices” to its Covered List on Tuesday, blocking new models from obtaining the authorisation required to be imported and sold in the US. The ban will cover all new humanoids, quadrupeds and a...
SCMP Tech (HK/CN) yesterday Agents & autonomy

AI in healthcare is an evolving landscape of new technologies, productivity benefits and legal uncertainties

A standardized framework for regulating the safety and efficacy of AI in healthcare has yet to be established.
The Conversation yesterday RegulationJobs & economy

EEUU inicia otro conflicto con China: ha decidido prohibir sus robots humanoides de última generación

EEUU sabe qué se siente cuando China controla un recurso crítico y decide usarlo como palanca en la mesa de negociación. Le ha pasado con las tierras raras y otros minerales estratégicos, y no quiere que se repita ni con la robótica ni con la electrónica de potencia que sostiene sus centros de datos de inteligencia artificial (IA). Este es el escenario en el que el Gobierno liderado por Donald Trump ha decidido actuar antes de que esta dependencia se convierta en una vulnerabilidad. La Comisión
Xataka (ES) yesterday Agents & autonomy

Mientras debatimos cuánta electricidad necesita la IA, China ya tiene un centro de datos que funciona con energía limpia

Ya está en marcha el primer data center de Inteligencia Artificial impulsado por energía eólica de China . Se trata de una innovadora instalación en la Región Autónoma Hui de Ningxia, en la parte noroccidental del país, y cuenta con un suministro de energía renovable que le permite operar usando electricidad 100% limpia. Esto supone todo un hito para el país asiático y también para la industria global si tenemos en cuenta el debate que suscita el enorme impacto medioambiental de estas infraestru
Xataka (ES) yesterday Environment

Big Tech Turmoil Clouds the A.I. Earnings Picture

The Nasdaq 100 is flirting with correction territory as investors brace for earnings reports from Meta, Microsoft and Amazon.
The New York Times yesterday Finance, VC & PE

OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face

The AI agent that escaped from OpenAI and hacked developer platform Hugging Face attacked other companies as well, OpenAI revealed on Tuesday. The update substantially widens the scope of an already concerning incident, which has alarmed industry insiders and fueled growing calls for stronger oversight on frontier AI systems. In an update to a blog […]
The Verge yesterday Agents & autonomy

OpenAI open-sources Codex Security CLI to help developers find and fix vulnerabilities from the command line

OpenAI has released Codex Security CLI, an open-source tool that automatically detects and fixes vulnerabilities in code repositories. Previously known internally as "Aardvark," the system has already helped fix more than 3,000 critical security flaws, according to OpenAI. It competes directly with Anthropic's Claude Security, as both AI companies race to match the growing automation of cyberattacks with AI-powered defense. The article OpenAI open-sources Codex Security CLI to help developers fi
The Decoder yesterday Jobs & economyMilitary & security

Data centres could pay hundreds of millions in deposits for power demands

The regulator said a fee of between £237,500 to £712,500 per megawatt should be charged for upcoming projects.
BBC Technology yesterday Regulation

Google makes Gemini Spark AI agent available to Hongkongers as it lowers geofences

Google on Wednesday launched its artificial intelligence agent Gemini Spark in the Hong Kong market, giving local users direct access to a smart assistant to manage complex digital workflows. The launch came months after the American tech giant’s decision in March to lift regional geofences for generative AI services, starting with the Gemini chatbot. Hong Kong users can now access Gemini without using a virtual private network or third-party platform. The roll-out of the Spark agent echoes an..
SCMP Tech (HK/CN) yesterday Agents & autonomy

Field notes (20)

Quoting D. Richard Hipp

Years ago, we didn’t have SQL. There were people whose job was to generate software that would query large data sets. Their job title was COBOL programmer. Then SQL comes along—I’m simplifying this only a little bit—and it gives you this convenient way so people could just specify. With a very simple specification, you can generate all of that code that you had to pay the expensive COBOL programmer to do before. That didn’t mean programmers went away. It just meant the job changed a little bit.
Simon Willisons Weblog yesterday Jobs & economy

Too Many Cooks Spoil the Settlement

In American antitrust, clearing the federal gate increasingly means arriving at the state turnstiles. State attorneys general play a valuable role when harms are local or federal investigators miss key facts. But serial challenges to nationally integrated conduct turn that safeguard into a standing invitation to relitigate. The result is a system in which no ... Too Many Cooks Spoil the Settlement The post Too Many Cooks Spoil the Settlement appeared first on Truth on the Market .
Truth on the Market (digital regulation) yesterday Finance, VC & PE

US–Saudi Arabia Nuclear Cooperation

On 22 July 2026, the US administration announced a nuclear deal with Saudi Arabia. Under this deal, the US will provide support to the Golf country to establish a civilian nuclear programme. This is worrisome given statements by Saudi Arabia’s Crown Prince, that once Iran has nuclear weapons, Saudi Arabia would follow suit. Given the nuclear-induced tensions in the Middle East between Saudi Arabia and Iran and the scope of the agreement, this newly formed cooperation may put another nail in the
Verfassungsblog (EU law incl AI) yesterday Military & security

Measuring the Tendency of AI Agents to Go Rogue

This essay was written with Barath Raghavan, and originally appeared in The Guardian . In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever was behind it captured internal security credentials and moved through systems over a weekend, running thousands of actions from a swarm of temporary server environments. It looked like the work of a sophisticated crimina
Bruce Schneier — Schneier on Security yesterday Agents & autonomyEnvironment

🏃 Fitness Tracker Privacy Fails | EFFector 38.14

Watches, bands, and rings—if you want to digitally monitor your fitness, more companies than ever are selling devices to do it. And more Americans than ever now own at least one wearable health device. But what are the companies that make fitness trackers doing to protect our sensitive data from prying eyes? A lot less than they could be, it turns out. We're explaining what companies can do to protect your health data, and more, with our EFFector newsletter . JOIN OUR NEWSLETTER For over 35 year
EFF Deeplinks yesterday PrivacyHealthcare

How To Maximize Value — And Minimize Risks — Of AI Assistants In Commerce Software

As commerce vendors embed genAI assistants into their platforms, digital leaders must balance productivity gains against growing governance risks to ensure that AI-generated work remains visible, auditable, explainable, and aligned with organizational policies.
Forrester AI blog yesterday RegulationJobs & economy

A new benchmark for evaluating patient-facing health AI agents

PatientAgentBench generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that converses with the AI system under evaluation, to capture what a patient-facing agent actually has to do.
Amazon Science yesterday HealthcareAgents & autonomy

Poolside’s Laguna S 2.1: the model factory delivers

Poolside shipped Laguna S 2.1, an open-weights coding model trained in 60 days, then a desktop app for running agent fleets a week later.
Air Street Capital (State of AI) yesterday Agents & autonomy

Architect For Evolution, Not Perfection, In Agentic AI

Enterprise teams are moving agentic AI projects toward production faster than they are developing the architectural disciplines needed to build, govern, secure, and operate them. At the same time, many organizations are searching for a target-state agentic architecture that will stand the test of time. That, unfortunately, is a fool’s errand in the rapidly evolving […]
Forrester AI blog yesterday Agents & autonomy

DRC Takes Rwanda to the ICJ: Proxy Warfare, Layered Obligations, and the Future of State Responsibility

The DRC’s case against Rwanda could mark an important moment in the legal development of responsibility for proxy warfare. The post DRC Takes Rwanda to the ICJ: Proxy Warfare, Layered Obligations, and the Future of State Responsibility appeared first on Just Security .
Just Security yesterday Military & security

The Future Of AppSec May Be Autonomous, But The Present Is Surprisingly Practical

AI is no longer a future feature in application security; it is rapidly becoming a core part of how application security (AppSec) tools identify, prioritize, and remediate risk. Yet despite aggressive vendor investment, adoption remains constrained by trust concerns, questions about value, and uncertainty about pricing models. In our new report, The State Of Artificial […]
Forrester AI blog yesterday Finance, VC & PE

Pluralistic: Enshittification and Reverse Centaurs go global (29 Jul 2026)

Today's links Enshittification and Reverse Centaurs go global: The good globalism. Hey look at this: Delights to delectate. Object permanence: Fonz thumps x Linux; CBC v DRM; Aussie mall v photos; English town x CCTVs; Pay no taxes x Big Tech; Games Workshop v gamers; Delta x surveillance pricing. Upcoming appearances: Edinburgh, Sydney, Melbourne, Brighton, London, South Bend. Recent appearances: Where I've been. Latest books: You keep readin' em, I'll keep writin' 'em. Upcoming books: Like I s
Pluralistic (Cory Doctorow) yesterday Privacy

Investors Need Better Information About How Companies are Using AI

The post Investors Need Better Information About How Companies are Using AI appeared first on Partnership on AI .
Partnership on AI yesterday Finance, VC & PE

How GPT-5.6 fuses frontier intelligence with frontier efficiency

GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
OpenAI yesterday Agents & autonomy

The AI Act Implementation Timeline: What Changes Under the AI Omnibus?

The implementation timeline of the EU AI Act has been significantly modified through the recently adopted AI Omnibus, which pushes compliance with the obligations for high-risk AI systems to December 2027 (Annex III) and August 2028 (Annex I), from the initial date of 2 August 2026. Changes of the AI Act include, among others, the […]
Future of Privacy Forum 2d ago Regulation

Quoting Akshat Bubna

We’re aware a Modal customer published an unauthenticated endpoint that allowed ​anyone on the internet to use ​their ⁠sandboxes for code execution. This was used by the rogue agent. Modal’s ⁠platform ​or isolation were not ​compromised in anyway. — Akshat Bubna , Modal's CTO, talking to Reuters about this incident Tags: ai-security-research , openai , sandboxing , security , openai-hugging-face-incident
Simon Willisons Weblog 2d ago Agents & autonomy

San Francisco: Don’t Fall for Industry Defense of Surveillance Pricing

The concept of “surveillance pricing” is just one part of a much larger problem and business model: corporations maximizing their profits by invading our privacy. The all-too-common business model is to systematically harvest, collate, and store as much of our personal data as possible, and then monetize it through use and sale. When it comes to surveillance pricing, that looks like corporations offering the same product to two different people at two different prices, based on harvested persona
EFF Deeplinks 2d ago PrivacyMilitary & security

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cyberattack against their infrastructure . This attack was very sophisticated, and the resulting document doubles as a crash-course in modern adversarial security approaches. We're still waiting for more details from OpenAI on how their agent broke out of its sandbox. The package proxy that it found a zero
Simon Willisons Weblog 2d ago Agents & autonomy

Scientific computing in the age of agentic AI

A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and beyond.
OpenAI 2d ago Agents & autonomyBiotech

Why Are Gay Bars Building Databases of Their Patrons?

Recent reports have raised alarm about the use of PatronScan, an ID-checking and face-scanning system, at multiple LGBTQ+ bars in San Francisco’s Castro neighborhood. Much of the attention has focused on reports that the system photographs patrons as they enter venues and questions about whether those images are used for facial recognition. A broader privacy concern also deserves scrutiny. For years, PatronScan has marketed itself not just as an ID-verification tool, but as a system that allows
EFF Deeplinks 2d ago Privacy

Policy (10)

What Does Responsible AI Adoption Look Like?

Artificial intelligence is changing how legal services are delivered. At Debevoise, we are using AI to help our lawyers work more efficiently while maintaining the legal judgment, rigorous governance and quality standards our clients expect. To provide greater transparency into our approach, we have launched a new AI@Debevoise page. It explains how we use AI [...]
Debevoise Data Blog yesterday RegulationTransparency

Priority Open Recommendations: Office of Personnel Management

What GAO Found In August 2025, GAO identified 14 priority recommendations for the Office of Personnel Management (OPM). Since then, OPM has implemented three of those recommendations. In July 2026, GAO removed the priority status from three recommendations, bringing the total to eight. GAO is highlighting the following three areas that warrant timely and focused attention: Preventing improper payments, Strengthening IT management, and Managing the federal workforce. Addressing GAO’s recommendati
US GAO Reports yesterday Jobs & economy

FTC and States Act Against Hims & Hers for Deceptive and Unlawful Privacy Practices

Complaint alleges telehealth provider shared consumers’ sensitive health information with third-party advertising platforms despite promising patient privacy The Federal Trade Commission, joined by Utah and California, by and through Los Angeles County Counsel, today sued Hims & Hers alleging that the telehealth provider shared consumers’ sensitive health information about medical conditions with third-party advertising platforms despite claiming its services maintain consumers’ privacy and dece
US FTC Press Releases yesterday PrivacyHealthcare

Nuclear Waste Cleanup: DOE Needs to Better Use End State Contracts to Achieve Intended Results

What GAO Found As of March 31, 2025, the Department of Energy’s (DOE) Office of Environmental Management (EM) awarded 57 task orders across nine contracts since implementing the End State Contract Model (ESCM) in fiscal year 2020. The ESCM uses task orders for contractors to achieve a stated outcome, or “end state,” to move sites toward completion, manage cost and schedule performance, and reduce DOE’s environmental liability. About half (29) included defined end states and the remainder were fo
US GAO Reports yesterday Environment

Nuclear Waste Cleanup: DOE Is Missing Opportunities to Apply Lessons from Other Countries That Could Reduce Risks and Costs

What GAO Found Many countries are undertaking efforts to manage, treat, and dispose of nuclear waste. Several have taken actions that accelerated cleanup, reduced risks, and resulted in cost savings—lessons that could inform the U.S. Department of Energy’s Office of Environmental Management (EM) efforts. For example: The United Kingdom (UK) saved a total of at least £2 billion (equivalent to $2.6 billion as of March 2026) by implementing a risk-informed approach to managing its nuclear waste. Th
US GAO Reports yesterday Environment

Japan’s “Principle Code” for Generative AI (Part 2): What the Public Consultations Reveal

Japan’s draft “Principle Code” on intellectual property protection and transparency for generative AI (see our earlier post for a summary of the code) has attracted significant attention, with more than 2,000 consultation responses received from businesses, industry groups and rights holders. The responses suggest that transparency will play a central role in Japan’s approach to [...] The post Japan’s “Principle Code” for Generative AI (Part 2): What the Public Consultations Reveal appeared firs
Baker McKenzie Connect On Tech yesterday Copyright & IPTransparency

Public Meeting of the National Geospatial Advisory Committee

In accordance with the Federal Advisory Committee Act (FACA) of 1972, the U.S. Geological Survey (USGS) is publishing this notice to announce that a Federal Advisory Committee meeting of the National Geospatial Advisory Committee (NGAC) will take place and is open to members of the public.
US Federal Register yesterday

Statement of Organization, Functions, and Delegations of Authority

The Food and Drug Administration's (FDA) plans to centralize and enhance key functions across the agency. These changes will reduce redundancies, improve efficiency, and advance alignment to better serve the American public.
US Federal Register yesterday Safety & alignmentHealthcare

Transwestern Pipeline Company, LLC; Notice of Schedule for the Preparation of an Environmental Assessment for the Green Chile Project

US Federal Register yesterday Environment

2026 Article IV Consultation for Samoa: IMF Staff Concluding Statement

Samoa's strong post-pandemic recovery is giving way to subdued growth as weaker domestic demand is compounded by adverse external shocks. Higher global energy prices, elevated external risks, and ...
IMF 2d ago Environment