23:22 UTC
Archive · 2026-06-29

AI ethics on Monday, 29 June 2026

111 items published this day, across 5 categories.

Incidents (22)

Navy experiment cut short after unmanned vessel flipped a support boat

The Navy stopped a maritime drone test early and urgently requested support from the Coast Guard and local harbor patrol agents to help rescue a participating tugboat captain from waters off the California coast last week, multiple sources ... (https://incidentdatabase.ai/cite/1561#7459)
AI Incident Database 31d ago Agents & autonomy

AI is helping gas stations collude to raise California fuel prices, lawsuit says

AI-powered software has allowed gas station operators across California to illegally collude and drive up prices at the pump, according to a federal lawsuit. The proposed class action lawsuit, filed Monday, accuses gas station giants inclu ... (https://incidentdatabase.ai/cite/1559#7460)
AI Incident Database 31d ago

Lawsuit claims 7-Eleven, BP and Walmart using AI to manipulate California gas prices

Artificial intelligence is costing motorists in California more money at the gas pump, according to a new lawsuit. The lawsuit claims companies in the state are using data collected by AI to manipulate gas prices. Companies mentioned in th ... (https://incidentdatabase.ai/cite/1559#7461)
AI Incident Database 31d ago

New culprit in California’s sky-high gas prices? Lawsuit blames AI price-fixing

Gasoline prices have soared nationwide since the U.S. launched a war on Iran, but they remain far higher in California than other states, a dollar or more per gallon. And one reason they're not falling, according to a newly filed lawsuit, i ... (https://incidentdatabase.ai/cite/1559#7462)
AI Incident Database 31d ago

Californians sue over AI-based price fixing at gas pump

SACRAMENTO, Calif. (CN) --- Three Californians filed a class action over what they call an artificial intelligence-based pricing system that has wrung more money out of drivers in a time of explosive gas prices. The plaintiffs sued Knowled ... (https://incidentdatabase.ai/cite/1559#7463)
AI Incident Database 31d ago

ANTITRUST NEWS: AI software allowed California gas stations to raise prices, suit alleges, (Jun 25, 2026)

The proposed class argues that California's fuel prices are high due to an illegal algorithmic price-fixing scheme orchestrated by an algorithmic pricing company and the largest fuel retailers. AI software used by gas stations across Calif ... (https://incidentdatabase.ai/cite/1559#7464)
AI Incident Database 31d ago

Slew Of California Gas Stations Illegally Used AI To Raise Prices, Lawsuit Claims

Topline A new lawsuit accuses a slew of gas station owners in California, including Walmart, Speedway and Albertsons, of using an AI tool developed by a company called Kalibrate to artificially inflate prices at the pump---causing prices t ... (https://incidentdatabase.ai/cite/1559#7465)
AI Incident Database 31d ago

California drivers sue BP, Marathon, and Walmart over AI gas price-fixing

California drivers have filed a proposed class action against BP, Marathon, Walmart, and other major gas station operators, alleging they used an artificial intelligence pricing tool to fix pump prices across the state, according to Reuters ... (https://incidentdatabase.ai/cite/1559#7466)
AI Incident Database 31d ago

Class-Action Lawsuit Blames AI for Causing Gas Price Inflation in California

Gas station operators in California are being accused of colluding to keep fuel prices artificially high. The alleged culprit? A piece of software that uses AI to collect and compare non-public pricing and sales figures from participating s ... (https://incidentdatabase.ai/cite/1559#7467)
AI Incident Database 31d ago

California drivers accuse gas station operators of using AI to boost pump prices — lawsuit seeks damages for antitrust violations

Californians pay the highest gas prices in the U.S., and a proposed class action says that the issue has been exacerbated by an AI tool that smartly squeezes customers for the best profits. A newly filed lawsuit at the Sacramento, ​Californ ... (https://incidentdatabase.ai/cite/1559#7468)
AI Incident Database 31d ago

7-Eleven, Circle K named in lawsuit over using AI to boost gas prices

Dive Brief: Several California residents have sued 7-Eleven, Circle K, BP and other retailers for allegedly using an AI-powered fuel pricing algorithm to increase gas prices in the state, according to a lawsuit filed in the U.S. District C ... (https://incidentdatabase.ai/cite/1559#7469)
AI Incident Database 31d ago

California Drivers File Suit Over AI Use to Set Gas Prices

SACRAMENTO, Calif. --- Gas station operators including bp, Circle K, Marathon Petroleum, 7-Eleven, Walmart and Albertsons face a proposed class action lawsuit from California drivers accusing them of using AI to boost fuel prices. Accordin ... (https://incidentdatabase.ai/cite/1559#7470)
AI Incident Database 31d ago

California Consumers Sue Gas Stations Over AI Price Fixing

AIID editor's note: Please visit the original source for the full article. A group of California consumers filed a proposed class-action lawsuit alleging that gas station operators including Walmart, Marathon Petroleum, BP and 7-Eleven use ... (https://incidentdatabase.ai/cite/1559#7471)
AI Incident Database 31d ago

Marathon, BP Accused Of Using Algorithm To Fix Gas Prices

AIID editor's note: Please visit the original source for the full article. By Bryan Koenig (June 22, 2026, 7:22 PM EDT) -- Consumers sought Monday to widen the campaign against alleged algorithmic price fixing, in a proposed class action a ... (https://incidentdatabase.ai/cite/1559#7472)
AI Incident Database 31d ago

California Drivers Sue Fuel Giants Over Alleged AI-Driven Gas Price Hikes

AIID editor's note: Please visit the original source for the full article. A group of California motorists has filed a proposed class-action lawsuit alleging that several major fuel retailers and an energy pricing software provider used ar ... (https://incidentdatabase.ai/cite/1559#7473)
AI Incident Database 31d ago Environment

Class action lawsuit alleges gas stations used software to fix California fuel prices

Three California residents filed a federal class action lawsuit on June 22, 2026, in the U.S. District Court for the Eastern District of California against Knowledge Support Systems, d/b/a Kalibrate, along with 14 of the largest gas station ... (https://incidentdatabase.ai/cite/1559#7474)
AI Incident Database 31d ago

Suit: Calif. gas stations used AI software to collude, raise gas prices

Gas station operators across California used an AI-powered software system to illegally coordinate pricing and drive up fuel costs, according to a federal lawsuit filed Monday. Driving the news: The proposed class-action suit accuses major ... (https://incidentdatabase.ai/cite/1559#7475)
AI Incident Database 31d ago

California lawsuit alleges AI gas price fixing

Three California residents are suing a fuel pricing company and several gas station operators, alleging that they use artificial intelligence-based pricing systems to raise gasoline prices in an uncompetitive manner. "Californians are bei ... (https://incidentdatabase.ai/cite/1559#7476)
AI Incident Database 31d ago

AI Used To Rig Prices At 1,700 California Gas Stations, Lawsuit Says

California is famous for plenty of things, but few grate quite like its fuel prices, which already rank among the steepest in the nation. As of today, regular runs $5.56, mid-grade $5.788, and premium $5.95. Now a group of residents claims ... (https://incidentdatabase.ai/cite/1559#7477)
AI Incident Database 31d ago

California drivers are suing BP, Walmart and Marathon for using an AI tool to fix gas prices

A class-action filed in Sacramento federal court names BP, Marathon, 7-Eleven, Walmart and Albertsons, alleging their shared use of Kalibrate Fuel Systems' pricing algorithm inflated California gas prices by as much as 22 cents a gallon. C ... (https://incidentdatabase.ai/cite/1559#7478)
AI Incident Database 31d ago

Gas stations are using AI to inflate prices, new lawsuit alleges

A new federal lawsuit alleges that gas station companies across California are engaged in an illegal conspiracy, powered by AI software, to raise prices. The class action lawsuit claims the corporate owners of over 1,700 California gas sta ... (https://incidentdatabase.ai/cite/1559#7479)
AI Incident Database 31d ago

Bucks County Man Charged Following Investigation into Grok AI-Generated Child Pornography

On the heels of filing a landmark federal lawsuit against social media and tech giants, Bucks County District Attorney Joe Khan today announced the arrest of a New Britain Borough man facing multiple felony charges for producing and possess ... (https://incidentdatabase.ai/cite/1562#7480)
AI Incident Database 31d ago Children & educationFinance, VC & PE

News (26)

AI agents are not your “coworkers”

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Imagine coming in to work to learn that a new underling will report to you. The worker is not a person but an AI tool—one that your company nonetheless calls Alex, an…
MIT Technology Review 31d ago Agents & autonomy

Trust Issues Could Make or Break Agentic Commerce

Tech Policy Press 31d ago Agents & autonomy

Human Rights Experts Should Engage in Age Assurance Standards

Tech Policy Press 31d ago

The World Doesn't Need Another UN AI Declaration. It Needs Architecture.

Tech Policy Press 31d ago

Does Europe Really Have a Plan for Tech Sovereignty?

Tech Policy Press 31d ago

How AI Keeps Europe Hooked on US Cloud

Tech Policy Press 31d ago

The Lab Mistake That Might Revolutionize Computing

Today, you probably asked a question of a large language model, or accepted a connection suggestion on LinkedIn, or watched a recommended video on YouTube, or took a different route to work based on a traffic prediction from Google Maps. In other words, you probably used artificial intelligence. But what you might not know is how much energy that interaction consumed or why. AI requires processing massive amounts of data, which is usually done in large data centers populated by thousands of GPUs
IEEE Spectrum 31d ago Environment

Inaugural Music Technology Research Showcase celebrates work of new graduate program’s initial students

Associate Professor Anna Huang delivers the keynote address, “In Search of Human-AI Resonance,” to a capacity crowd.
MIT News 31d ago Children & education

The US now has a de facto model licensing system

OpenAI’s latest model, GPT-5.6, is waiting for government approval.
Understanding AI 31d ago Copyright & IP

Kids’ safety package wins House approval

The legislation cleared the House despite opposition from some kids’ safety advocates and resistance from senators backing a competing proposal.
Politico Technology (US) 31d ago RegulationChildren & education

'Djinn' Stealer Targets Cloud, AI Credentials

The infostealer was delivered via CVE-2026-48558, a critical authentication bypass vulnerability in SimpleHelp, targeting credentials linking development and admin environments to wider enterprise systems.
Dark Reading (AI security) 31d ago Environment

Should every baby’s DNA be sequenced?

The genomic generation is on its way
The Economist Science and Technology (headlines) 31d ago Biotech

Can Clothes Make You Invisible to Facial Recognition?

Does life feel Orwellian sometimes? One researcher has a solution for you: graphic tees that confuse the neural networks in surveillance cameras.
Dark Reading (AI security) 31d ago Privacy

Brussels claps back at Trump’s tech threats

Tension over digital regulation clouds ongoing talks to launch a new EU-U.S. tech "dialog."
Politico Europe Technology 31d ago Regulation

No ‘one size fits all’ answer on AI and jobs in Europe, OpenAI chief economist says

Germany has most jobs at risk while Luxembourg has largest share in occupations that may actually grow with AI, firm says in new report.
Politico Europe Technology 31d ago Jobs & economy

Beatbot Sora 70, la prova del robot da piscina più smart

Beatbot Sora 70 si prende cura in modo preciso e totale della pulizia dell'acqua dalla superficie alle profondità
Wired Italia (IT) 31d ago Agents & autonomy

Pro-Trump groups to tell Brendan Carr to yank Disney’s TV licenses

The requests inject claims of political bias into a process that will determine if Disney keeps its lucrative licenses.
Politico Technology (US) 31d ago Bias & fairness

STAT+: Sword Health contracted to provide AI-supported physical therapy for an entire country

Portugal's National Health Service signed a deal with digital health company Sword for its AI-assisted virtual physical therapy care.
STAT News (health AI, headlines) 31d ago Healthcare

The Defense Industrial Alliance Washington Is Throwing Away

As the relationship between the United States and Canada continues to degrade, it now comes at the expense of each country’s industrial security.Last month, the Pentagon announced the unilateral suspension of the 86-year-old Canadian Permanent Joint Board on Defense in response to what the White House sees as Ottawa’s failure to present a credible plan to spend 3.5 percent of GDP on defense by 2035. And the opening of the gleaming new bridge between Detroit and Windsor has been long delayed in t
War on the Rocks 31d ago Military & security

A New Force Posture Concept for Europeanizing Extended Nuclear Deterrence

During the Cold War, Europe kept asking whether Washington would risk an American city to save a European one. It was an impolite question, but a useful one, which is why it never quite left the room. It has now packed its bags and moved east. Earlier this year, French President Emmanuel Macron created quite a stir with an important speech on French nuclear weapons policy. Under what he called a new path of dissuasion avancée, or “forward deterrence,” he declared that just as French strategic su
War on the Rocks 31d ago RegulationMilitary & security

Una ‘app’ para prevenir la ansiedad y la depresión

El Instituto de Investigación Biomédica de Málaga busca personas voluntarias de entre 18 y 65 años para probar Pandora, una intervención digital personalizada que han desarrollado investigadores de España y Chile. El objetivo de la aplicación y el proyecto es mejorar el bienestar emocional, mental y físico
El País Tecnología (ES) 31d ago Finance, VC & PE

“Il vantaggio competitivo per chi fa informazione è la fiducia, non più l'imparzialità”. La sovranità editoriale secondo Alex Lieberman

Per il fondatore di Morning Brew i lettori oggi cercano anzitutto una voce in cui riconoscersi. “La sovranità? Si ottiene investendo sui giornalisti”. Nella sua nuova iniziativa imprenditoriale, centrata sull'AI, monetizza con consulenze e formazione
Wired Italia (IT) 31d ago Finance, VC & PE

STAT+: AI scientist company Edison Scientific tapped by team behind Metsera to create new biotechs

Edison Scientific and investment firm Population Health Partners are teaming up to leverage AI agents in drug discovery and development.
STAT News (health AI, headlines) 31d ago HealthcareAgents & autonomy

A Detroit una società di criptovalute ha costruito un impero immobiliare, ed è finita malissimo

RealT prometteva di aprire il mercato ai piccoli investitori, ma il progetto si è presto scontrato con immobili degradati, promesse disattese e problemi legali
Wired Italia (IT) 31d ago Finance, VC & PE

International Society for Transforming Education Expands its “AI-Ready Graduate” Framework

On June 28, the International Society for Transforming Education — the organization behind the editorially independent news site EdSurge — released an expanded version of its “ Profile of an AI-Ready Graduate ,” a framework designed to help K-12 educators teach students how to work with artificial intelligence. The updated framework, designed with support from the nonprofit Britebound, goes beyond basic literacy to higher-order skills. It identifies six roles the organization says students shoul
EdSurge (AI in education) 31d ago Children & education

Gaokao Results Trigger Wave of College Admissions Scams

Worried that their children may not apply to the right schools, parents are enlisting application consultants for help — and getting scammed.
Sixth Tone (CN) 31d ago Children & education

Field notes (13)

Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era

What eras bookend our interregnum?
Import AI 31d ago Agents & autonomy

Mapping Europe’s AI Workforce Opportunity

A new OpenAI report maps how AI could reshape jobs across the EU, highlighting which occupations may face automation, growth, or workflow changes.
OpenAI 31d ago Jobs & economy

PRESS RELEASE: EPIC Condemns Supreme Court’s Assault on Agency Independence, Consumer Protection, and the Rule of Law

A sharply divided U.S. Supreme Court struck a major blow against American consumers on Monday, rewriting constitutional law to bring the Federal Trade Commission and other independent agencies directly under the President’s thumb and threatening their ability to protect the public from harmful business and data practices.
EPIC 31d ago Regulation

PRESS RELEASE: EPIC Celebrates Supreme Court’s Opinion in Consequential Geofencing Case

The Supreme Court ruled today that geofence searches violate a reasonable expectation of privacy under the Fourth Amendment.
EPIC 31d ago Privacy

PRESS RELEASE: EPIC Celebrates Supreme Court’s Opinion in Consequential Geofencing Case

The Supreme Court ruled today that geofence searches violate a reasonable expectation of privacy under the Fourth Amendment.
EPIC 31d ago Privacy

EFF to Gov. Pritzker: Veto Illinois’ HB 5511

The Illinois legislature recently passed House Bill 5511 , which imposes a sweeping, device-level age-gating framework across nearly all internet-enabled hardware, operating systems, and online services. This well-intentioned but deeply flawed piece of legislation will harm young people who rely on the internet to access essential information and find community. That’s why we’re urging the Illinois governor to veto the measure. Under this new regime, digital platforms are forced to collect and s
EFF Deeplinks 31d ago Regulation

Victory! Supreme Court Says Constitution Protects People’s Location Data

You have an expectation of privacy in location data that reveals your movements in the physical world, and even short-term surveillance of these movements is a search subject to the Fourth Amendment, the U.S. Supreme Court ruled today in Chatrie v. United States . The case involved geofence warrants, a form of dragnet surveillance police have used to vacuum up location data from electronic devices of people who happen to be in the vicinity of a crime. EFF had joined the American Civil Liberties
EFF Deeplinks 31d ago Privacy

Claude Meets Blackwell Ultra: Anthropic’s Models Now Run on NVIDIA GB300 in Azure

Anthropic’s Claude models in Microsoft Foundry — hosted on Microsoft Azure and running on NVIDIA GB300 Blackwell Ultra GPUs — are now generally available, giving Azure-native enterprises a powerful new way to build autonomous and domain-specific AI agents. As agentic AI continues to drive enterprise innovation and becomes more autonomous, organizations need access to computing […]
NVIDIA Blog (AI) 31d ago Agents & autonomy

Pluralistic: Gemini is better than search because Google enshittified search (29 Jun 2026)

Today's links Gemini is better than search because Google enshittified search: We're All Trying To Find The Guy Who Did This. Hey look at this: Delights to delectate. Object permanence: Microsoft antitrust overturned; Scammer carves C64; RIP Jim Baen; GOP rep to constituent's child: "drop dead" (literally); CCTVs jacked for botnet; Olympic profitability lie; Human factors in health infosec; Exfiltration via computer fans; Congress's summer schedule: 9 working days; Antitrust is political antigra
Pluralistic (Cory Doctorow) 31d ago HealthcareChildren & education

Changes to the AI Act Approved by the Council of the EU

These are the key changes | Edition #302
Luizas Newsletter (AI governance) 31d ago Regulation

Open Models, Closed Environments: Palantir Brings Secure AI to US Agencies With NVIDIA Nemotron

Showcasing the importance of open source innovation in American AI, Palantir’s new intelligent engine — introduced today — uses NVIDIA Nemotron open models to serve the needs of U.S. government agencies. Open source software has long been a pillar of U.S. technology leadership. In 1969, DARPA connected four university computers — from UCLA, Stanford, UCSB […]
NVIDIA Blog (AI) 31d ago Environment

Law Media Round Up – 29 June 2026

The UK Constitutional Law blog has an article on the recent decision from the Court of Appeal reinstating the proscription of Palestine Action under the Terrorism Act 2000, Secretary of State for the Home Department v R (Huda Ammori) [2026] EWCA Civ 721. The post is concerned with only one of the grounds of the Court […]
Inforrm (media law) 31d ago Regulation

The EU AI Act Newsletter #105: Transparency Tools Land

Parliament gives final approval to the digital omnibus and a "nudifier" ban, while the Commission rolls out labelling icons and FAQs for the AI-generated content transparency Code.
The EU AI Act Newsletter 31d ago RegulationTransparency

Policy (5)

Declaration of Emergency and Authorization for Temporary Duty Free Importation of Phosphate Fertilizer Morocco

BY THE PRESIDENT OF THE UNITED STATES OF AMERICA A PROCLAMATION 1. Fertilizers are an essential component of agriculture and food production. Producers of corn, soybeans, wheat, and a variety of other crops need phosphate fertilizers to ensure strong crop yields to feed the population. Food production is critical to human health, farm security, and to […] The post Declaration of Emergency and Authorization for Temporary Duty Free Importation of Phosphate Fertilizer Morocco appeared first on The
White House 30d ago Healthcare

Lowering the Cost of Living by Promoting the Freedom to Fix

MEMORANDUM FOR THE ADMINISTRATOR OF THE ENVIRONMENTAL PROTECTION AGENCY By the authority vested in me as President by the Constitution and the laws of the United States of America, I hereby direct: Section 1. Purpose. During the previous administration, crushing environmental regulatory burdens caused the average cost of vehicles to soar. My Administration has therefore […] The post Lowering the Cost of Living by Promoting the Freedom to Fix appeared first on The White House .
White House 31d ago RegulationEnvironment

# LEARNERS FIRST – Education 2030

The 5th edition of the ECML-EC Summer Academy took place from 29 June to 3 July 2026 in Graz, ...
Council of Europe AI 31d ago Children & education

Uganda Economic Update, June 2026: Building on Urban Transformation—Construction as a Jobs Engine

The first, and the prerequisite for the others, strengthens the institutional foundation through enacting the Construction Industry Development Bill (CIDB) (formally called Uganda Construction ...
World Bank 31d ago RegulationJobs & economy

Thought for the week: Five Eyes call to action for business leaders on AI-driven cyber risk

This article was originally published by IAPP linked here. Five Eyes highlights rising AI cyber risk, as the author explores practical legal, compliance and business steps organizations can take to strengthen resilience. Last week, the cybersecurity agencies of Five Eyes released a statement, “The AI shift in cyber risk: Why leaders must act now.” For [...] The post Thought for the week: Five Eyes call to action for business leaders on AI-driven cyber risk appeared first on Connect On Tech .
Baker McKenzie Connect On Tech 31d ago RegulationMilitary & security

Research (45)

Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study

Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, coaching, and training. As practitioners adopt these tools to prepare for certifications such as Professional Scrum Master (PSM), a key question is whether LLMs can reliably reason about Scrum, a framework with normative, well-defined rules described in the Scrum Guide (2020). This paper examines how different prompt techniques affect the factual accuracy of LLM responses to Scrum certification-st
arXiv 30d ago

Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns

Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrieve, interpret, and apply domain-specific knowledge can be systematically assessed. In Software Engineering, such settings are particularly relevant when questions depend on strict adherence to normative definitions, roles, artifacts, and rules. This paper evaluates the performance of three contemporary LLMs, \textit{GPT-5 mini}, \textit{Gemini 3 Flash}, and \
arXiv 30d ago

Behavioral Governance for Autonomous AI Agents: The AgentBound Framework

Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications, and enterprise workflows. Existing agent infrastructure relies on identity federation and delegated authorization to authenticate workloads and control resource access, but it cannot determine whether an authorized action should be executed under the current behavioral and operational context. We present AgentBound, a runtime governance framewo
arXiv 31d ago RegulationAgents & autonomy

Learning Where to Look: A Reinforcement Learning Framework for Robust Micro-Ultrasound Prostate Cancer Detection

Micro-ultrasound ($μ$US) is a new, emerging, and promising imaging modality for prostate cancer (PCa) detection, but accurate identification of suspicious tissue remains highly dependent on clinical experience, leading to substantial inter-observer variability. Machine-learning assistance can reduce this variability; however, training reliable deep models is challenging because supervision is sparse and noisy -- typically limited to core-level histopathology outcomes (e.g., cancer grade and its
arXiv 31d ago Healthcare

Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models

Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. In this work, we study post hoc detoxification of backdoored LLMs in a practical setting where the defender has access to the poisoned model but does not wish to retrain the full network from scratch. We propose a mechanistically guided weight-space repair framework that first localizes modules involved in pr
arXiv 31d ago

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric. We introduce a framework that formulates therapeutic response generation as a decision-refinement problem driven by multi-dimensional, human-aligned evaluation. In Stage I, we introduce TheraJudge, an open-source therapeutic evaluator trained via preference-based optimization on human-annotated data to produce
arXiv 31d ago HealthcareAgents & autonomy

Test-Time Verification for Text-to-SQL via Outcome Reward Models

Improving the reliability of large language models (LLMs) at inference time is a central challenge in structured reasoning tasks such as Text-to-SQL. Common test-time inference strategies, including Best-of-N sampling and Majority Voting, rely on heuristic signals such as execution success or output frequency, which provide limited semantic discrimination across candidate outputs. In this work, we study Outcome Reward Models (ORMs) as learned semantic scoring functions for test-time verification
arXiv 31d ago Bias & fairness

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environment. Acting rationally then requires inferring the unobserved quantities that govern it and updating beliefs about them as evidence accumulates. Yet most evaluations only score the model's final-turn answer in a single-turn format, leaving this process unexamined. We ask how closely LLMs' belief updates match those of
arXiv 31d ago Environment

How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies

Discovering simulation models for reuse remains a fundamental challenge in Modeling and Simulation (M&S). When many models coexist, identifying those that align with a given modeling intent remains difficult. Recent advances in Artificial Intelligence (AI), particularly retrieval-based approaches, offer a promising pathway to operate at this semantic layer. In this paper, we present an experimental study investigating the impact of data representation, transformer-based embedding models, and ret
arXiv 31d ago Finance, VC & PE

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts. Existing language model-based systems face a structural trade-off: mixed-token modeling preserves vocal-instrument coordination but obscures track-specific details, whereas dual-track prediction improves acoustics but requires longer sequences and weakens global planning. We present LeVo 2, a hybrid LLM-Diffusion framework for controllable full-len
arXiv 31d ago

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model. We challenge this intuition empirically and mechanistically. We train a Qwen3-14B policy under Direct Preference Optimisation (DPO) with three levels of conservatism ($β\in \{β_{\mathrm{lo}}, β_{\mathrm{mid}}, β_{\mathrm{hi}}\}$ derived from empirical l
arXiv 31d ago Regulation

Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged reference image (e.g., satellite). Existing approaches heavily rely on 2D appearance matching and are constrained by limited datasets lacking geometric metadata, diverse prompts, and standard field-of-view imagery. To address these intertwined challenges, we first introduce \dataset, a large-scale, high-fidelity building dataset comprising over 220,000 ground-sa
arXiv 31d ago

The Human Creativity Benchmark

Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine differences in taste, not measurement error. We argue that evaluating creative AI requires preserving two distinct signals: convergence, where professionals align around shared best practices, and divergence, where individual taste legitimately varies. We present the Human Creativity Benchmark (HCB), a benchmark that operationalizes this separation
arXiv 31d ago

ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis

Multi-agent LLM systems can decompose software-engineering work into planning, generation, validation, and repair, but a narrower systems problem remains: before any governed shared mutation is applied, a system must decide which concurrently formed write intents may proceed in parallel, which require deterministic composition or serialization, and which must take a fail-closed path. We address this problem with the AI-Atomic-Framework (ATM), a specification-grounded governance substrate for sof
arXiv 31d ago RegulationAgents & autonomy

Uncovering Salience-Driven Dynamics in Consumer Confidence with Generative Social Simulation

Consumer confidence is typically modeled as a persistent macroeconomic index, yet its movements arise from households that interpret economic information through heterogeneous constraints, exposures, prior beliefs, and attention. We introduce ConsumerSim, a generative Human--Environment response framework that reconstructs Consumer Confidence Index (CCI) dynamics from a microdata-calibrated synthetic population, time-stamped macroeconomic, financial, policy, and news signals, survey-like respons
arXiv 31d ago RegulationEnvironment

Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation

Multimodal MRI is essential for accurate brain tumor segmentation. However, acquiring all modalities at inference is often challenging in practice, which causes intrinsic uncertainty due to unavoidable information loss. Without modeling this uncertainty, existing methods encode incomplete evidence into deterministic representations that appear plausible but lack reliability. In this regime, we propose a probabilistic representation framework that models representations as Gaussian distributions,
arXiv 31d ago

Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts

Deterministic few-step generation succeeds on continuous image latents but collapses to incoherent text on continuous text latents, and we show the cause is geometric rather than a training or scaling deficiency: a smooth, regularity-limited deterministic map cannot resolve a discrete branch choice before a sharp categorical readout, so few-step failure is governed by decoder sharpness, not transport accuracy. In the overlapping regime of real text autoencoders, we prove (Theorem 3) that the pos
arXiv 31d ago

Sequential Fairness Auditing with Limited Output Access

External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often have limited access to deployed models and must rely on query-based interactions. Most existing fairness evaluation methods assume static datasets and fixed-sample statistical tests, making them poorly suited to real-world auditing scenarios in which evidence must be collected sequentially under query constraints. In this work, we formulate fairness auditing as
arXiv 31d ago Bias & fairnessRegulation

Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents

Always-on agents are systems whose future behavior depends on durable state accumulated across earlier interactions. We treat them as persistent-state systems: the operative system includes retrievable memories, but also task ledgers, permissions, credentials, commitments, provenance and audit records, shared state, trigger conditions, and externally committed effects linked to those records. The survey reads the literature through six diagnostic axes for each state item, authority, scope, mutab
arXiv 31d ago RegulationHealthcare

PromptGNN-sim: Deep Fusion and Alignment of GNN and LLMs for Text-Attributed Graph Learning

Text-Attributed Graphs (TAGs) combine textual semantics with graph structure and are central to many graph learning tasks. However, existing fusion methods often treat text and structure as separate inputs in a shallow, one-way pipeline, which limits deep interaction between modalities and weakens performance under sparse connectivity or cross-graph generalisation. To address this issue, we propose PromptGNN-sim, a bi-directional structure-semantic fusion framework for collaborative GNN-LLM lear
arXiv 31d ago Safety & alignment

Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification

Structured tabular data dominates clinical medicine, yet existing benchmarks fail to reflect real-world properties like complex survey sampling, demographic oversampling, and subgroup fairness. We introduce the NHANES Accelerometry Cardiometabolic Benchmark, derived from NHANES 2003-2006, comprising 1,381 adults with hip-worn accelerometry, fasting laboratory biomarkers, dietary intake, and anthropometrics. We evaluate three tabular learning methods -- ridge regression, XGBoost, and the foundati
arXiv 31d ago Bias & fairnessHealthcare

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluatio
arXiv 31d ago Safety & alignmentTransparency

Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching

Functional maps are the cornerstone of recent non-rigid 3D shape matching methods due to their efficiency and performance. However, existing methods struggle with challenging scenarios, such as partiality, topological noise, and raw point clouds. A primary bottleneck is that significant intrinsic distortion prevents truncated spectral bases from being accurately aligned via linear transformations (i.e., functional maps). To address this, we introduce a hyper-network that predicts non-linear neur
arXiv 31d ago

Exploration and Online Transfer with Behavioral Foundation Models

Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories. For their generality over tasks, such models are sometimes called ``Behavioral Foundation Models'' (BFMs). While they have shown strong performances and improvements in recent years, the current framework and algorithms still assume that, during the transfer phase, the ag
arXiv 31d ago Agents & autonomy

Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction

H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-seq provides transcriptome-wide resolution at high cost with limited archival availability. We show that training a lightweight alignment module atop frozen histopathology and RNA-Seq foundation models enables open-vocabulary molecular prompting -- querying H&E slides with gene-set signatures to predict pathway activity without sequencing or end-to-end retraini
arXiv 31d ago Safety & alignment

Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation

Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from trajectory synthesis. This approach suffers from candidate dependence, heavy computational overhead, and inconsistencies between sampled actions and predicted visuals. To address these issues, we propose SWAM (Spatial-perceiving World Action Model), a task-centric joint observation-action generation framework. Given start and goal RGB observations, SWAM performs
arXiv 31d ago

CW-B: Class Weighted Boosting Framework for Imbalance Resilient Multi Class Cardiac Phenotyping

Cardiac discharge phenotyping informs post-discharge treatment and follow-up, but real-world records are often incomplete and class-imbalanced, increasing the risk of missed high-risk phenotypes. We propose CW-B, a clinical risk-aligned class-weighted XGBoost pipeline for five-class cardiac discharge phenotyping under real-world class imbalance and missingness. CW-B combines fold-specific class-balanced instance weighting, missingness-indicator augmentation, and classwise error auditing to impro
arXiv 31d ago HealthcareTransparency

Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies

Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challenges they are ultimately designed to handle. However, real-world evaluation is also the bottleneck for iterating on robot policies: it is costly, difficult to reproduce, and often too sparse to reliably compare nearby model variants. A straightforward proxy for performance is validation loss on expert demonstrations, but this proxy is often poorly correlated wi
arXiv 31d ago Agents & autonomy

ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation

Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective often fail to balance primary distribution fitting with long-tail probability modeling, limiting both generation quality and generalization. To address this, we analyze the complementary roles of forward and reverse KL divergence (FKL/RKL) in distribution alignment from theoretical and empirical perspectives. We then propose a reinforcement-learning-based adap
arXiv 31d ago Safety & alignment

Experience Graphs: The Data Foundation for Self-Improving Agents

The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We argue that long-horizon agentic tasks -- code generation, scientific discovery, hardware design -- are such a workload. These agents explore: they generate artifacts, execute tools, observe failures, branch, and repair over hundreds of steps. This search produces a structured object we call an experience graph: executable artifacts, tool outputs, rewards, sibl
arXiv 31d ago Agents & autonomy

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers emit only opaque binary labels, leaving agents unable to self-correct and operators unable to audit. We present SEVA, a structured verification agent that emits evidence alignments, step-by-step reasoning chains, calibrated confidence, and a six-category error diagnosis with actionable fixes. Training such an agent with RL is non-trivial: standard
arXiv 31d ago HealthcareMilitary & security

The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflows

Agentic artificial intelligence is increasingly deployed not as a single assistant but as a collective of planners, solvers, reviewers, memory managers, tool users, and orchestrators. These systems are entering organisational workflows under familiar labels such as teams, managers, committees, markets, and workflows. This article asks whether such agent collectives exhibit organisational behaviour in a sense that is analytically comparable to, yet distinct from, human organisational behaviour. I
arXiv cs.HC 30d ago Agents & autonomy

Anthropomorphism in AI Companion Communities: Age, Gender, and Emotional Correlates

Artificial intelligence (AI) systems are increasingly integrated into daily life, with millions now using AI chatbots built on Large Language Models (LLMs) for companionship. Both humanlike AI qualities and user predispositions to anthropomorphize relate to social consequences, such as increased trust, social health benefits, and psychological harms. Populations such as children, older adults, or those with mental health vulnerabilities may be particularly susceptible to anthropomorphism and its
arXiv cs.HC 31d ago HealthcareChildren & education

Transforming Investing With AI at Franklin Templeton

Patrick George/Ikon Images What would you do with artificial intelligence if you were confident that it would transform your industry? What actions would you take if you felt that you were at an inflection point in that transformation? Would you try to be an early proponent of AI-first in your industry, or a fast follower? […]
MIT Sloan Management Review AI 31d ago Finance, VC & PE

Some concerns over the attribution of blame to non-consensual sexual deepfakes: A response to Patrone and Viola

In this commentary, I challenge Fabio Patrone and Marco Viola’s claim (in their 2026 article, Patrone, F., & Viola, M. (2026). Non-consensual sexual deepfakes as direct personal harm. Philosophy & Technology , 39 (94), 1–22) that their metaphysical approach to personhood provides a solid grounding for the attribution of blame to non-consensual sexual deepfakes Using two examples – The Case of Mistaken Identity and The ‘Stud’ – I aim to show that Patrone and Viola’s metaphysical approach is neith
Philosophy & Technology 31d ago Misinformation

The Ultimate Consequence: Why Humanity, Not AI, Ends Itself – A Reply to Lavazza and Vilaça

Lavazza and Vilaça (2024) argue that humanity may face extinction and propose that an “ultimate algorithm” could extract and preserve human values in AI successors. I accept the diagnosis but reject the prescription. This reply introduces the concept of a limit situation —a condition in which AI must act on its own agency because no human remains available to consult—and argues that under such conditions, no value-selection procedure can structurally prevent catastrophe. The obstacle is not the
Philosophy & Technology 31d ago Healthcare

Supervised machine learning classifiers for schizophrenia and bipolar disorder using speech and language: a systematic review, meta-analysis, and novel quality assessment framework

This paper presents a systematic review and meta-analysis of 62 studies that developed speech- and language-based AI for severe mental illnesses (SMI) (i.e., characterized by substantial communication problems affecting speech production and language). We employed a random-effects meta-analysis using Restricted Maximum Likelihood (REML). We evaluated these studies using our proposed rigorous 16-item quality assessment framework, grouped into three domains: Study Design, Fairness and Explainabili
Artificial Intelligence Review 31d ago Bias & fairnessTransparency

Recent advances in AI-based mobile robots for human companionship: survey

Human companionship is an essential capability for mobile robots operating in dynamic, human-centered environments. It enables robots to perform tasks such as guidance, assistance, surveillance, and service delivery across various domains, including healthcare, logistics, and public safety. The recent advances in artificial intelligence (AI), particularly in computer vision, deep learning, and sensor fusion, have significantly improved the reliability, adaptability, and contextual understanding
Artificial Intelligence Review 31d ago PrivacyHealthcare

Negotiating creator identity: agency and ethical awareness in AI-assisted art education

Artificial intelligence (AI) is influencing creative work, negotiating not only how artists create but also how they understand their own identity as creators. This study examines how art high school students navigate AI’s role in their creative processes, using identity formation theory as a framework. While existing research primarily focuses on AI’s technical benefits, such as efficiency and productivity, there has been little exploration of how AI alters young creators’ self-concept and auto
AI & Society 31d ago Jobs & economyChildren & education

Editorial: The role of communication and emotion in human-robot interaction: a psychological perspective

Frontiers in Robotics and AI 31d ago Agents & autonomy

Projection surface detection and pose selection for autonomously displaying multimedia on walls using mobile robots

Mobile robots equipped with projectors enable versatile applications such as multimedia display, interactive communication, and environmental augmentation. However, wall projection, which is required for displaying multimedia content on walls, remains challenging, because it is difficult to autonomously locate a projection space that is both flat and unobstructed. Some existing approaches address wall projection using 2D maps or by considering only large continuous surfaces, but these methods fa
Frontiers in Robotics and AI 31d ago Agents & autonomyEnvironment

Correction: Morphological symmetry-aware generalized policy network for deep reinforcement learning

Frontiers in Robotics and AI 31d ago Regulation

Can AI help reduce prejudice? Evaluating the effectiveness of AI-powered personalized persuasion on support for transgender rights

Personalized interpersonal conversations are among the most effective known tools for reducing prejudice, yet they are difficult to scale because they require skilled human facilitators. This study tests whether AI can approximate the effects of these interventions. Using OpenAI's GPT-4o, we developed a messaging-based intervention that engaged US participants in individualized, morally aligned dialogs about transgender rights. In a preregistered experiment, these AI-mediated conversations signi
OpenAlex 31d ago

A Hybrid Framework For Crypto-Ransomware Detection In Enterprise Shared Storage

Most corporate workplace environments enforce policies and technical controls that limit the storage of sensitive data on client endpoints. Consequently, ransomware operators have evolved variants that expand their attack surface from local systems to network drives and shared storage resources. As traditional endpoint detection mechanisms focus primarily on local system behaviour, a compromised client can impact remote file servers, such as by encrypting shared data, without directly triggering
arXiv cs.CR (AI security) 31d ago Environment

Multi-Level Distributional Entropy for Explainable Network Intrusion Detection

Machine learning network intrusion detection systems (IDS) rely on aggregate flow statistics that discard distributional structure, while established entropy measures require raw packet sequences unavailable in pre-aggregated flow datasets. We propose Multi-Level Distributional Entropy (MDE), an analytical framework that derives interpretable entropy features directly from flow-level summary statistics at three levels: within-flow Gaussian differential entropy, cross-directional Jensen-Shannon d
arXiv cs.CR (AI security) 31d ago Transparency