23:23 UTC
Archive · 2026-06-24

AI ethics on Wednesday, 24 June 2026

73 items published this day, across 4 categories.

Incidents (11)

Belfast man found not criminally responsible for killing wife, attacking mother in Readfield

AUGUSTA --- A Belfast man who killed his wife and assaulted his mother with a fire poker in Readfield was deemed not criminally responsible for those acts Thursday. He was ordered held by the state in a psychiatric hospital over the objecti ... (https://incidentdatabase.ai/cite/1491#7440)
AI Incident Database 36d ago

Scammers use AI photo of missing dog at emergency vet to steal nearly $2,000 from St. Pete couple

ST. PETERSBURG, Fla. --- A couple in St. Petersburg is scammed out of nearly two thousand dollars in bogus vet bills. They say they fell for it because the scammers took pictures of a lost dog posted online and sent the couple an AI-genera ... (https://incidentdatabase.ai/cite/1477#7441)
AI Incident Database 36d ago

AI Detection Was Built for Faces. Climate Deception Targets Environments.

Much of the public debate around generative AI has focused on human-centered content: political deepfakes, manipulated speeches, fabricated celebrity videos, and non-consensual intimate imagery. This focus is justified given the scale, seve ... (report_number: 7442)
AI Incident Database 36d ago MisinformationEnvironment

AI system fails during Glendale Community College graduation ceremony

GLENDALE, AZ (AZFamily) --- An AI system used to read graduate names at Glendale Community College's commencement ceremony malfunctioned, leaving students and families frustrated. The names being read during GCC's commencement didn't appea ... (https://incidentdatabase.ai/cite/1503#7443)
AI Incident Database 36d ago Children & education

Sixth Circuit Sanctions Attorneys for Fake Citations – What Does This Mean for Use of AI?

The Sixth Circuit finally weighed in on the use of fake cases hallucinated by artificial intelligence. A panel recently sanctioned two Tennessee attorneys for a smorgasbord of misconduct during merits briefing, including citing fake cases. ... (https://incidentdatabase.ai/cite/1447#7444)
AI Incident Database 36d ago

Fabricated Citations, Real Consequences: The Sixth Circuit’s Multi-Pronged Response to Attorney Misconduct

The Sixth Circuit's recent decision in Whiting v. City of Athens, 2026 U.S. App. LEXIS 7479, is a significant ruling addressing attorney misconduct, particularly the submission of fabricated legal citations and misrepresentations in appella ... (https://incidentdatabase.ai/cite/1447#7445)
AI Incident Database 36d ago

#FactCheck: Viral AI Video Showing Finance Minister of India endorsing an investment platform offering high returns.

Executive Summary: ‍ A video circulating on social media falsely claims that India's Finance Minister, Smt. Nirmala Sitharaman, has endorsed an investment platform promising unusually high returns. Upon investigation, it was confirmed tha ... (https://incidentdatabase.ai/cite/1553#7446)
AI Incident Database 36d ago Finance, VC & PE

Fifth Circuit Sanctions Lawyer for AI-Generated Brief

The Fifth Circuit recently issued a sharp warning to lawyers who use generative AI in federal appellate practice. In Fletcher v. Experian Information Solutions, Inc. (opinion below), the court sanctioned appellate counsel after finding that ... (https://incidentdatabase.ai/cite/1453#7447)
AI Incident Database 36d ago

Fifth Circuit Sanctions Opinion Gives Practical Advice for AI Use

In Fletcher v. Experian Information Solutions, Inc., the U.S. Court of Appeals for the Fifth Circuit sanctioned a lawyer $2,500 for filing a reply brief with several "hallucinated" case citations and then providing evasive responses to the ... (https://incidentdatabase.ai/cite/1453#7448)
AI Incident Database 36d ago

Fletcher v. Experian Info Solutions: Appellate Sanctions for AI-Hallucinated Briefing and Lack of Candor Under FRAP 46(c) and Inherent Authority

1\. Introduction Fletcher v. Experian Info Solutions is an appellate sanctions decision arising from a Fair Credit Reporting Act dispute in which the Fifth Circuit confronted a rapidly recurring litigation risk: the use of generative artif ... (https://incidentdatabase.ai/cite/1453#7449)
AI Incident Database 36d ago

South Africa withdraws AI policy due to fake AI-generated sources

JOHANNESBURG, April 27 (Reuters) - South Africa has withdrawn its first draft national AI policy after revelations that it ​contained fictitious sources in its reference list ‌which appeared to have been AI-generated. "The most plausible e ... (https://incidentdatabase.ai/cite/1467#7450)
AI Incident Database 36d ago Regulation

News (19)

Anthropic's Red Lines Are No Substitute for Public Law

Tech Policy Press 36d ago Regulation

EU Lawmakers Press Commission on Child Safety as Debate on Age Limit Heats Up

Tech Policy Press 36d ago Children & education

A Policy Playbook to Inoculate the Public Against AI Text Falsehoods

Tech Policy Press 36d ago Regulation

The EU's AI Transparency Code of Practice, Explained

Tech Policy Press 36d ago Transparency

Congress Should Pass AI Law to Reassure the Public

Tech Policy Press 36d ago Regulation

What Alex Bores’ defeat tells us about AI politics

The NY-12 House race drew over $27m from various AI PACs. But it’s hard to unpick their impact.
Transformer 36d ago

Nuclear Stability in the Age of AI

In 2024, Paul Scharre and Michael Depp wrote, “Artificial Intelligence and Nuclear Stability,” where they argued integrating artificial intelligence into the nuclear chain of command presents both opportunities and risks. Two years later, as AI becomes increasingly integrated into military systems and processes, we asked them to revisit their arguments. Image: Senior Airman Jason Wiese via Wikimedia CommonsIn 2024, you argued that integrating artificial intelligence into the nuclear chain of com
War on the Rocks 36d ago Military & security

STAT+: A dispatch on AI from BIOtech’s big summer bash

In this edition of STAT's AI Prognosis: Brittany Trang brings the latest from BIO on how biotech companies are approaching artificial intelligence.
STAT News (health AI, headlines) 36d ago Biotech

STAT+: AI wades into a vexing medical mystery: What causes sudden cardiac death?

A new study published in Nature uses artificial intelligence to identify people at high risk for sudden cardiac death, and pinpoints a possible reason.
STAT News (health AI, headlines) 36d ago Healthcare

Rhino Horn, Leopard Skin and Tiger Claws Sold Openly on Facebook

Warning: Includes graphic descriptions of animal harm and images of animal parts from the outset. A Bellingcat investigation has uncovered a Myanmar-based wildlife trafficker who has operated openly across social media for at least six years, claiming to have sold tiger bones, rhino horn, elephant skin and other products from protected and endangered species to […] The post Rhino Horn, Leopard Skin and Tiger Claws Sold Openly on Facebook appeared first on bellingcat .
Bellingcat (tech investigations) 36d ago Finance, VC & PE

How governments enable kleptocrats by doing nothing

I was listening to Ezra Klein interviewing a left-wing Democratic Party strategist the other day about what a post-Trump U.S. foreign policy might look like, and it was a pretty striking demonstration of Europe’s irrelevance right now that the only Western European country mentioned in the 90 minutes of the chat was the UK, and The post How governments enable kleptocrats by doing nothing appeared first on Coda Story .
Coda Story (authoritarian tech) 36d ago Regulation

Designing Drones for Africa

This exclusive Cogs of War interview is with Maxwell Maduka, the co-founder and chief engineer of Terra Industries, an African defense technology company building autonomous drone and counter-drone systems designed for the continent’s operating conditions. As cheap imported airframes flood African markets and non-state actors employ drones across the Sahel, we asked Max why Terra is betting on Africa building its own defense industrial base.Cheap Turkish and Chinese drones and sensor systems hav
War on the Rocks 36d ago Military & security

AI Agents and the Unseen Work of War

Armies run on more than what happens at the front. Behind every operation is a vast amount of coordination, administration, logistics, and judgment. Bill Pessin, senior vice president of national security at Salesforce and a former U.S. Army logistics officer, joins Jonathan to discuss how military organizations can use AI agents, what makes these tools different from ordinary software, and why safety and accountability matter when new technology enters national security work. They also discuss
War on the Rocks 36d ago RegulationMilitary & security

Advocacy Groups Express Mixed Views on Embryo Editing

At least two new start-up companies, Preventive and Origin Genomics, say they are developing strategies that combine gene editing with in vitro fertilization to correct disease-causing mutations. Advocacy groups for people living with these genetic disorders have been quiet on the developments.
Undark Magazine 36d ago Biotech

Alibaba reportedly seeks sale of gaming unit Lingxi Games, valuation starts at $1.03 billion

Alibaba Group is planning to sell its gaming unit Lingxi Games, according to people familiar with the matter. Alibaba has approached at least five potential buyers, including Chinese game developers 37 Interactive Entertainment, China Ruyi, Century Huatong and Giant Network, as well as two private equity firms, the sources said. The business is being marketed […]
TechNode (CN) 36d ago Bias & fairnessFinance, VC & PE

Podcast: Who Is Really in Charge When Tech Enters the Classroom?

Two educators are reckoning with who is really in charge: technology or the teacher.
EdSurge (AI in education) 36d ago Children & education

Outgrowing the Chromebook: Why Advanced STEM Demands Better Student Tech

#ASUSEducation @ASUS
EdSurge (AI in education) 36d ago Children & education

Vibe Coding Sparked a Love of Reading in My Classroom

Lessons learned from a year of building an AI literacy tool.
EdSurge (AI in education) 36d ago Children & education

Student Sues Chinese Airline After 10-Minute Flight Change

The 19-year-old said he believed the flight change policy was unfair to customers, who have to bear the burden of changes in departure times, while airlines face no cost.
Sixth Tone (CN) 36d ago RegulationChildren & education

Field notes (8)

How Algorithmic Systems Govern Kenya’s Content Moderators

An exclusive survey of AI workers in Kenya reveals how automated management affects their livelihoods. Unions and advocacy groups are beginning to fight back.
AlgorithmWatch 36d ago

FPF’s 2026 DC Privacy Forum: Leading Voices in AI, Privacy and Emerging Technology

By Paige Garvin, FPF Communications Intern The Future of Privacy Forum hosted its third annual DC Privacy Forum: Advancing Principled Data Protection, AI, and Digital Governance Practices on June 10th, 2026. This year’s Forum gathered government officials, academics, civil society representatives, and privacy professionals to discuss developments in AI governance, privacy regulation, youth online safety, […]
Future of Privacy Forum 36d ago RegulationPrivacy

🦅 Domestic Spying Takes an L | EFFector 38.12

Sold to the public as a foreign surveillance tool, Section 702 is the law has let intelligence agencies spy on millions of Americans’ private conversations without a warrant. Despite years of revelations about this law's misuse, Congress has repeatedly reauthorized Section 702 without meaningful reform. Until this month, that is, when it finally lapsed in a major victory for privacy. In our latest EFFector newsletter , we're covering the expiration of Section 702 and what happens next . JOIN OUR
EFF Deeplinks 36d ago RegulationPrivacy

Trump’s Iran War: The Midwife To A Renewable Energy Future

The post Trump’s Iran War: The Midwife To A Renewable Energy Future appeared first on NOEMA .
Noema Magazine 36d ago Environment

Will we fix AI bias against LGBTQ+ users?

A new report shows how AI systems are already failing LGBTQ+ users. The problem is: it may also the best way to fix moderation issues that traditional systems never managed to address.
Everything in Moderation (Ben Whitelaw) 36d ago Bias & fairness

The opposite of America's AI problem is happening in Brazil

While the US debates whether to regulate AI at all, Brazil has built the most detailed AI-and-elections rulebook of any democracy and the gap between the two is becoming a headache for companies
Anchor Change (Katie Harbath) 36d ago Regulation

The CEO of AWS on why Amazon is hiring 11,000 interns and junior employees

Matt Garman argues that junior employees are as necessary as ever. But AWS now sells agents that can recruit, code, and process claims. Will the balance hold?
Platformer 36d ago Agents & autonomy

Scaling Laws, Carefully

Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, dataset size $D$, and compute $C$, following a power-law curve, which appears as a straight line on a log-log plot. We can view scaling laws as a framework for describing the relationship between compute, loss, model size and data; at its core, it is about how to allocate precious compute optimally between $N$
LilLog (Lilian Weng) 36d ago Regulation

Research (35)

MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation

Retrieval-augmented generation (RAG) over knowledge graphs has emerged as a promising approach for grounding large language models, yet existing benchmarks largely overlook the challenges of retrieval in multimodal knowledge graph RAG (MKG-RAG). In practice, retrieval is a critical bottleneck: multimodal knowledge is heterogeneous, difficult to align across modalities, and often poorly served by retrievers designed for unstructured corpora. To address this gap, we introduce MKG-RAG-Bench, a cros
arXiv 35d ago

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image, offering no way to evaluate reasoning over observed human behavior. We introduce WatchAct, a benchmark for robot manipulation grounded in observed human behavior. Each instance pairs a real-world human-a
arXiv 36d ago Agents & autonomy

Scoring Is Not Enough: Addressing Gaps in Utility-fairness Trade-offs for Ranking

Scoring functions are used to represent the relevance of individual documents. In modern information retrieval or recommendation systems, they are often learned from data and play a pivotal role in ranking sets of documents or items in a way that maximizes utility to a query or user. With the recent interest in algorithmic fairness, the success of scoring has naturally led to methods that learn scores that simultaneously trade off fairness and utility. In this work, we show that in stark contras
arXiv 36d ago Bias & fairness

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stake in the outcome) and uncertainty suppression (no explicit unknowns or hedges before committing to an action). We introduce narration-of-thought (NoT), a system prompt that structures chain-of-thought into five sections: protagonist, stakeholders, two-step consequences, uncertainty, then commitment. NoT adds no training, parameters, or fine-tuning. On 100 Dai
arXiv 36d ago

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-domain evaluations remain largely limited to static knowledge recall. This is a critical gap for a sector that requires live data retrieval, specialized regulatory and market knowledge, and multi-step quantitative reasoning under real-world constraints. We present an empirical study of tool-augmented LLM agents on real-world energy market analytics t
arXiv 36d ago RegulationHealthcare

Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems

Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment. This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning but by requiring independently attested evidence at the point of consequential action. We formalise this institutional pattern as a computational governance model for AI agent systems. Under the proposed model, an agent retains full auton
arXiv 36d ago RegulationHealthcare

Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System

Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess. However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the granular quality of gameplay. Nevertheless, incorporating move-by-move information into rating adjustments presents a significant challenge given the substantial noise and the vastness of the game-state space. To address this, we propose the Drift-Diffusion-Enhanced Elo Rating System (DD-Elo
arXiv 36d ago

Learning Action Priors for Cross-embodiment Robot Manipulation

Most Vision-Language-Action (VLA) models build on a Vision-Language Model (VLM) backbone by attaching an action module and optimizing the full policy jointly. This design inherits strong visual and linguistic priors from the VLM, but leaves the action module to learn physical motion almost from scratch. As a result, the policy lacks an explicit motion prior, forcing early optimization to simultaneously discover temporal action dynamics and cross-modal alignment, a challenge further amplified in
arXiv 36d ago RegulationSafety & alignment

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems

AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places controls inside the agent's own runtime: system prompts, output filters, and guardrail libraries. Any control in the agent's address space is reachable by inputs that influence it; this generalizes to any AI system with sufficient reach into its own runtime, a class we term escapable AI systems. We identify four properties that an authorization mecha
arXiv 36d ago Safety & alignmentAgents & autonomy

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining

Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ("Sue cried because"), it resolves the next pronoun to she, generalizing to held-out probes (0.94 by step 925). By step 3,500 the same model scores near zero on the same probes, although the rule's evidence is still in the training data. We call this within-run reversal natural ungrokking: the corpus decides, with no trace in the loss curve, which learned rules a model keeps
arXiv 36d ago

Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis to study socio-technical power structures at scale. We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (co
arXiv 36d ago RegulationAgents & autonomy

Statistical and Structural Approaches to Algorithmic Fairness

Modern machine learning systems have outgrown their origins as isolated predictive constructs, evolving into complex socio-technical architectures that actively mediate human opportunity. As algorithms increasingly determine access to economic and social opportunities, it has become widely recognized that these systems are deeply embedded with the structural inequalities and prejudices of their environments. The field of algorithmic fairness emerged in response to the growing recognition that mo
arXiv 36d ago Bias & fairnessEnvironment

SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

As multimodal conversational systems increasingly engage in spoken interaction, their ability to navigate paralinguistic social cues has become a critical bottleneck for natural human-AI communication. However, existing evaluations of machine emotional intelligence assess reasoning exclusively through isolated text or passive acoustic perception, overlooking the complex cross-modal reasoning required for active, multi-turn dialogue. We introduce \textsc{SpeechEQ}, a comprehensive framework desig
arXiv 36d ago

Feedback-Coupled Memory Systems in Continuous Time

The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which - the agent update operator $f_i$ and the environmental update operator $Ψ$ - are left axiomatically undefined in the original framework. To address this, $f_i$ is defined by Mechanism-Based Intelligence (MBI), where agents update locally through a decentralized price mechanism and economic principles, and $Ψ$ is defined by the Coupled Memory Graph Process (CM
arXiv 36d ago Agents & autonomyEnvironment

AI Snitches Get Glitches: Towards Evading Agentic Surveillance

To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many employers (and even nation-states) already provide their users with this technology. However, widespread adoption of AI agents creates a new risk to abuse access to user data for another goal: surveilling users. These users might not even have the ability or permission to control the actions and data accesses of the surveilling agents. We introduce and f
arXiv 36d ago PrivacyAgents & autonomy

Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction

A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requirement specifications, with minimal or no manual redesign. In this setting, policy generation by AI agents can be a credible path when paired with a plant-aware validator (e.g., a digital twin) that can check generated candidate actions before execution. However, practical deployment is constrained by inference latency and compute footprint: large cloud-based mod
arXiv 36d ago RegulationAgents & autonomy

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment

Sparse Mixture-of-Experts (MoE) architectures have emerged as an increasingly influential paradigm as they offer a strategic balance between parameter scalability and computational efficiency. However, low-resource languages, which suffer from a scarcity of high-quality training data, often have their tokens routed to different experts than those predominantly activated by high-resource inputs, which limits cross-lingual expert sharing. This cross-lingual routing divergence consequently hinders
arXiv 36d ago Safety & alignment

Enterprise Data Asset Quality: A Management-Standard Conformity-Benefit Realization Framework and Formation Mechanisms

Motivated by the limited standardization of enterprise data asset quality evaluation and the unclear relationship between assessment outcomes and value realization, this study develops a three-dimensional framework comprising Data Asset Management Capability, Data Quality Standard Conformity, and Data Asset Benefit Realization Capability, based on grounded theory and LDA topic modeling. To examine the formation mechanisms of data asset quality, this study adopts a multi-method approach combining
arXiv 36d ago

Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa

Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat compute primarily as a technical input rather than as an outcome of investment, ownership, and financial control. This paper examines AI infrastructure investment flows across Africa through a systematic analysis of 46 publicly announced projects totalling USD $12.7 billion between 2019 and 2025. Using a value chain framework, we analyze who invests in AI-rel
arXiv 36d ago RegulationFinance, VC & PE

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation

With the widespread adoption of large language models (LLMs) in chatbots and everyday applications, companies increasingly need guardrails that are effective while remaining low-cost and low-latency. Safety evaluation of LLM outputs has generally relied on LLM-based judges, which can be effective but are often slow and expensive to deploy at scale. In this paper, we evaluate whether fine-tuned modern encoder classifiers from the ModernBERT family, including ModernBERT and Ettin, can reliably ide
arXiv 36d ago

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction

While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pathways. We present KG-TRACE, a novel neuro-symbolic framework that integrates the WHO mutation knowledge graph (KG) as a structured biological constraint on a neural genomic model. Unlike existing methods that learn statistical patterns in isolation, KG-TRACE fuses genomic features and RotatE-based KG embeddings through a learned epistemic trust gat
arXiv 36d ago Biotech

Position Spaces and Graphs

In this paper, we introduce position graphs, a graph-based reasoning framework based on the formalization of position spaces. This framework utilizes two strict partial orders, representing horizontal and vertical alignment and precedence, to model the relative positions of discrete tokens. Unlike general qualitative spatial calculi, position graphs are constrained by a chain condition and compatibility requirements that focus on rows and columns. We provide a comprehensive theoretical analysis
arXiv 36d ago Safety & alignment

Steering Vision-Language Models with Joint Sparse Autoencoders

Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representations that are difficult to use as controllable cross-modal steering directions. We introduce the Joint Sparse Autoencoder (JSAE), which uses an explicit alignment constraint to jointly factorize sequence-pooled vision and language activations into shared, interpretable image/caption-level features. Applied to LLaVA, JSAE recovers cross-modal feat
arXiv 36d ago Safety & alignment

LCG: Long-Context Consistent Image Generation with Sparse Relational Attention

Recent image generation models achieve impressive quality in single-image synthesis, but often fail to maintain consistency across sequential outputs, as required in comics, storyboards, and visual narratives. We propose Long-Context Generation (LCG), a framework for long-context multi-image text-to-image generation, to improve consistency and scalability in long-context multi-image generation. LCG employs the Sparse Relational Attention (SRA) mechanism to selectively attend to core features acr
arXiv 36d ago

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue), and nonverbal vocalizations (NVs). Moreover, collecting cross-lingual target speech that is both translation-faithful and expressively aligned with the source is difficult at scale, making reference-based evaluation impractical. We introduce STEB (Speech-to-Speech Translation Expressiveness Benchmark), a 32.6-hour Chin
arXiv 36d ago

The impact of artificial intelligence on enterprise software user roles

Artificial Intelligence (AI) is rapidly reshaping the nature of work in software development, transforming user roles, workflows, and collaboration patterns across enterprise platforms. This qualitative study investigates how AI alters professional responsibilities within the context of SAP's Business Technology Platform (BTP), combining expert interviews (n=20) and a participatory workshop (n=24). The results reveal substantial shifts in day-to-day tasks and roles in the development domain, cha
arXiv 36d ago Finance, VC & PE

TopoCast: A Topological Fidelity Framework for Evaluating Transformer-Based Time Series Forecasting

Deep learning-based models have achieved state-of-the-art performance in Time Series Forecasting (TSF), yet their evaluation remains dominated by pointwise error metrics such as Mean Squared Error (MSE), which quantify numerical accuracy but overlook structural properties of the forecast signal, including recurrent dynamics, oscillatory behavior, and phase alignment. As a result, forecasts exhibiting over-smoothing, phase shifts, or frequency distortions may achieve favorable error scores despit
arXiv 36d ago Safety & alignment

Conformal Recovery-Deadline Certificates for Runtime Assurance of Adapting Controllers

Runtime assurance (RTA) protects a safety-critical system by switching from an advanced controller to a verified safe controller when a monitored condition is violated. The standard latching rule, which trips on the first breach of the safe set and then coasts, is correct for a diverging controller but pathological for a capable online-adapting one. Such a controller is unsafe by design during a bounded recovery transient. It must excite the plant to identify the fault before it can correct it,
arXiv 36d ago

Agentic Knowledge Tracing: A Multi-Agent LLM Architecture for Stealth Assessment of Financial Literacy in Serious Games

Assessing financial literacy during gameplay without disrupting the learning experience remains a key challenge in serious games for education. We present the Agentic BKT pipeline, a multi-agent large language model architecture for stealth assessment of financial competencies from open-ended gameplay events. The pipeline processes events from a 2D platformer serious game aligned with the OECD/INFE financial literacy framework through four phases: (1) the game captures every player decision as a
arXiv 36d ago Children & educationAgents & autonomy

AI Coaching for Accelerating Human Skill Development with Reinforcement Learning

AI copilots can substantially boost human performance through shared control, but excessive assistance can induce over-reliance and skill atrophy. This paper studies how an embodied AI agent can act as a coach that accelerates human motor-skill development. We argue that effective coaching requires strategic scaffolding and stepping back that are aligned with the learner's capability, allowing productive failures that drive learning. We formalize the interactive AI coaching process as a non-coop
arXiv 36d ago Agents & autonomy

Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing

Large Language Models (LLMs) have shown promise for automated penetration testing, yet existing end-to-end black-box evaluations are highly susceptible to error cascading: failures in early reconnaissance can mask an agent's actual ability to exploit vulnerabilities. To more accurately characterize these capabilities, we propose a two-stage decoupled evaluation framework that separates exploit execution from reconnaissance. Using ground-truth injection and knowledge-driven ablation across 70 hig
arXiv 36d ago Agents & autonomy

Cross-Subject Predictive Validity for Learning Outcomes of Delayed Start Behavior

Behavioral detectors provide valuable insights into learner motivation and self-regulation. Among these, delayed start, a new session-level detector, has shown great promise as a valid behavioral measure that generalizes well across systems. In this paper, we examine cross-subject predictive validity of delayed start behavior. Using iReady data from 711 grade 7 students, we find delayed starts during Math practice are predictive of standardized test performance in both Math ($β$=.07 SD, p=.02) a
arXiv 36d ago RegulationChildren & education

Communicability-Inspired Positional Encoding (CIPE)

Positional encodings (PEs) are essential for Transformers. Yet designing effective PEs for non-Euclidean graphs remains challenging. Such encodings should ideally induce an Attention-Compatible Geometry for self-attention: not merely describing graph structure, but defining a geometry whose inner products reflect meaningful structural relatedness. To realize this geometry, we propose Communicability-Inspired Positional Encoding (CIPE), built from communicability, a measure between pairs of nodes
arXiv 36d ago

Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning

Training-time data poisoning during fine-tuning poses a significant threat to large language models (LLMs) deployed for abstractive text summarization, where small task-specific datasets exert disproportionate influence on model behavior. In this setting, adversaries manipulate fine-tuning data to induce persistent summarization failures, such as biased or harmful summaries, while preserving standard evaluation metrics. We present a unified post-hoc defense framework for detecting and remediatin
arXiv cs.CR (AI security) 36d ago Bias & fairnessMilitary & security

Can Machine Learning Break Wi-Fi Privacy? A Study on MAC Address Randomization

Medium Access Control (MAC) address randomization has been widely adopted during the IEEE 802.11 network discovery phase as a countermeasure against passive tracking. This paper exposes vulnerabilities in these privacy protocols by demonstrating that devices remain identifiable using Machine Learning (ML)-based fingerprinting. To study the potential tracking capabilities of a passive attacker, we evaluate different eavesdropping scenarios and configurations. To this end, we extract unencrypted h
arXiv cs.CR (AI security) 36d ago Privacy