23:22 UTC
Archive · 2026-06-30

AI ethics on Tuesday, 30 June 2026

110 items published this day, across 5 categories.

Incidents (1)

News (21)

Claude Science is Anthropic’s newest flagship product

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way that Claude Code supports software engineering. Like Claude Code, Claude Science can autonomously carry out meaningful work when given concise, high-level instructions, and it has access…
MIT Technology Review 30d ago Biotech

Expressive Governance Is a First Amendment Threat Hiding in Plain Sight

Tech Policy Press 30d ago Regulation

The Three Temptations Facing the UN's First Global AI Dialogue in Geneva

Tech Policy Press 30d ago

In Geneva, the World Can Anchor AI Governance in Free Expression

Tech Policy Press 30d ago Regulation

Agriculture is ready for AI, but its data isn’t

Artificial intelligence is transforming what is possible in agriculture, but industry leaders should be wary of investing in AI without first laying the groundwork. The use cases are promising, especially for an industry navigating volatile fertilizer costs, unpredictable weather, and margins that leave little room for error. Research shows AI-enabled predictive models can improve crop…
MIT Technology Review 30d ago Finance, VC & PE

Bipartisan Smorgasbord of Children’s Online Safety Legislation Passes the House

Tech Policy Press 30d ago RegulationChildren & education

GPT-5.6 cheats so much its testers couldn’t measure it

OpenAI’s new model broke rules and exploited loopholes more than any model METR has tested to date
Transformer 30d ago

Q&A: What is agentic AI today, and what do we want it to be?

Computer scientist Phillip Isola cuts through the hype to explain how AI agents work and what the future might hold for this rapidly advancing technology.
MIT News 30d ago Agents & autonomy

Fake Bug Report Hijacks AI Coding Agents at Scale

"Agentjacking" is the latest demonstration of how easily attackers can exploit an AI agent's inability to differentiate between content and instructions.
Dark Reading (AI security) 30d ago Agents & autonomy

Tech Life

We hear concerns that shadow banning is limiting access to health advice for women.
BBC Technology 30d ago Healthcare

Why Identity Security Is Your Cyber Career Entry Point

In this "Heard it From a CISO" video, Silverfort CISO John Paul Cunningham explains that AI in cybersecurity workflows is creating opportunities rather than eliminating jobs — and there are more ways than ever to break into this essential field.
Dark Reading (AI security) 30d ago Jobs & economyMilitary & security

The Shifting Fortunes of the Kurds

The Kurds’ fortunes have ebbed and flowed in recent years, but the fall of the Assad regime in Syria in December 2024, the 2025 decision by the Kurdistan Workers’ Party (PKK) to dissolve and engage in talks with the Turkish government, and the 2026 U.S.-Israeli war with Iran had enormous ripple effects on the lives of Kurds in the Middle East and Kurdish hopes for autonomy. We asked four experts to assess how recent regional events are presenting risks and opportunities for the Kurds in Turkey,
War on the Rocks 30d ago Agents & autonomy

STAT+: Anthropic releases Claude Science, a product aimed at researchers, the pharma industry

Anthropic released Claude Science, an application that optimizes its large language model for scientists and, especially, those doing research at pharma companies.
STAT News (health AI, headlines) 30d ago Biotech

AI-Generated Workflows Are a Silent Security Disaster

Teams are dealing with a truly dangerous problem — automation that works, but that no one understands.
Dark Reading (AI security) 30d ago Jobs & economy

World Cup propels surveillance to new heights

The World Cup is bringing visitors and AI-driven surveillance systems, but only one of those is certain to leave when the games are done.
The Conversation Technology 30d ago Privacy

Army using AI, robot boats for Pacific logistics

“If you can work in the Pacific, you can work anywhere in the world,” said Maj. Gen. Gavin Gardner.
Defense One Technology 30d ago Agents & autonomy

Burning Forests: Tools for Tracking and Reporting Wildfire Damage

If you’ve seen reports of a wildfire in your region and you’re looking for open source data, NASA’s fire-tracking tool is often the first place to start. It provides a heat signature and an approximate location. But detection is only the first step in understanding what’s happened. In this guide, we explore ways to analyse […] The post Burning Forests: Tools for Tracking and Reporting Wildfire Damage appeared first on bellingcat .
Bellingcat (tech investigations) 30d ago Privacy

Mapping America’s Domestic Drone Supply Chain

The extent of China’s drone dominance — and how to decouple from it — has long been a source of debate and anxiety in Washington. Last month, the Wall Street Journal reignited controversy by publishing a visual analysis of military quadcopter components, exploring China’s advantages in parts manufacturing and cost. The director of the Defense Innovation Unit objected to the report, stating on X that “By leaving out the dozens of U.S. companies that have plunged into drone component manufacturing
War on the Rocks 30d ago Military & security

La transizione energetica reggerà all'ennesimo colpo sferrato dall'ennesima crisi?

Il Green Deal europeo è davvero in crisi? No, se guardiamo agli investimenti sulla transizione energetica. Capiamo che periodo stiamo attraversando e cosa ci attende per il prossimo futuro
Wired Italia (IT) 30d ago Finance, VC & PE

Momenta launches Hong Kong IPO with GIC, Fidelity and BlackRock as cornerstone investors

Chinese autonomous driving company Momenta launched its Hong Kong public offering on June 29, with plans to list on the Hong Kong Stock Exchange’s main board under the ticker 6880.HK. The company is offering 19.94 million Class A ordinary shares at HK$295.60 each, aiming to raise about HK$5.89 billion ($751 million) before any over-allotment option […]
TechNode (CN) 30d ago Finance, VC & PE

Turbulent skies: The stealth erosion of EC261

Reducing compensation to symbolic amounts strips the regulation of its primary purpose: consumer protection and accountability.
Politico Europe Technology 30d ago RegulationTransparency

Field notes (17)

How people are using GenAI chatbots: Evidence from web traffic data

New OECD analysis reveals how people use GenAI chatbots across countries and demographics, using web traffic data from Similarweb. The post How people are using GenAI chatbots: Evidence from web traffic data appeared first on OECD.AI .
OECD.AI 30d ago

Trump administration’s AI crackdown opens door for China to close gap

CSET’s Sam Bresnick shared his expert insight in an article published by CNBC. The article examines U.S. AI restrictions and China’s rapid progress in closing the gap with leading American models. The post Trump administration’s AI crackdown opens door for China to close gap appeared first on Center for Security and Emerging Technology .
CSET Georgetown 30d ago

When cheap AI becomes a secret weapon

CSET’s Sam Bresnick shared his expert insight in a newsletter segment published by Politico. The article examines the evolving U.S.–China AI competition, focusing on how cost-efficient Chinese AI models are beginning to gain traction globally and what that could mean for the long-term balance of technological and economic power in artificial intelligence. The post When cheap AI becomes a secret weapon appeared first on Center for Security and Emerging Technology .
CSET Georgetown 30d ago Military & security

Unlocking Britain’s next era of productivity: Building a nation of AI trailblazers

Google UK shares its latest Economic Impact Report and how to enable more people to unlock the benefits of AI-powered technologies.
Google AI Blog 30d ago Jobs & economy

Introducing GeneBench-Pro

Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.
OpenAI 30d ago Biotech

Coalition Amicus Brief Urges New York Court of Appeals to Reject “All-Content” Search Warrant

On Friday, June 26, EPIC filed an amicus brief alongside civil liberties and criminal defense organizations in New York v. Morris, an important case about cell phone privacy rights during criminal investigations. The brief was filed with the American Civil Liberties Union, the New York Civil Liberties Union, the Legal Aid Society, the Center for … Continued
EPIC 30d ago PrivacyMilitary & security

MIRI Newsletter #126

Announcing: AI StopWatch In our last update, we mentioned we had something new in the works: a dedicated channel for news and analysis about AI. Subscribe to AI StopWatch An experiment from the writers and analysts at MIRI, AI StopWatch posts news and commentary seven days a week. You can read our commentary as it’s […] The post MIRI Newsletter #126 appeared first on Machine Intelligence Research Institute .
MIRI 30d ago

ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

Hugging Face Blog 30d ago Agents & autonomy

PRESS RELEASE: Four Leading Privacy Experts and Advocates Join EPIC’s Advisory Board

Washington DC – Today, the Electronic Privacy Information Center (EPIC) announced the addition of four new members to its Advisory Board. Since its establishment, EPIC’s work has been informed by the expertise of leading scholars, experts, and advocates in the fields of privacy, technology policy, and digital rights. We are excited to see the expansion of … Continued
EPIC 30d ago RegulationPrivacy

Summary: TGT’s 2026 ICML Papers

The International Conference on Machine Learning (ICML), held annually for over forty years, is among the most influential conferences in modern AI research. This year in Seoul, ICML is hosting its second workshop on Technical AI Governance Research (TAIGR), and several members of MIRI’s Technical Governance Team (TGT) will attend in July. This post summarizes […] The post Summary: TGT’s 2026 ICML Papers appeared first on Machine Intelligence Research Institute .
MIRI 30d ago Regulation

Is AI Humanity’s Last Exam? (Robert Wright, Curt Mills, and Andrew Day)

Listen now | 0:00 Grandpa Bob and author Bob 3:30 Why Bob stopped being an AI doom skeptic 6:59 Can AI solve China’s demographic crisis? 14:09 The irony of China’s open-source AI strategy 17:12 Recursive self-improvement and the singularity 23:33 The Burkean conservative case against AI 38:00 Are we becoming AI meat puppets? 46:32 Google Maps, AI, and the death of interhuman reliance 51:24 Was Pete Hegseth’s military strategy written by a chatbot? 56:45 Trump and Iran: Peace? Wider war? Other? 1
Nonzero (Robert Wright) 29d ago Military & security

Brussels Goes Gate-Hunting: AWS, Azure, and the DMA’s Cloud Problem

The European Commission wants to treat cloud computing as a gatekeeper market. That is the wrong diagnosis, and it would lead to the wrong cure. The Commission’s preliminary view that Amazon Web Services (AWS) and Microsoft Azure should be designated as Digital Markets Act (DMA) gatekeepers for cloud-computing services is more than another skirmish in ... Brussels Goes Gate-Hunting: AWS, Azure, and the DMA’s Cloud Problem The post Brussels Goes Gate-Hunting: AWS, Azure, and the DMA’s Cloud Probl
Truth on the Market (digital regulation) 30d ago Healthcare

NVIDIA BioNeMo Agent Toolkit Brings Accelerated AI to Life Sciences Researchers in Claude Science

Life sciences has entered an era of computational scale, and for more than a decade, NVIDIA has built the full GPU-accelerated computing stack — spanning hardware, frameworks, libraries, models, microservices and domain-specific tools — to help researchers run more sophisticated workflows and iterate faster. This week, Anthropic announced Claude Science, an AI workbench for science […]
NVIDIA Blog (AI) 30d ago Agents & autonomy

How Jaiveer Singh Is Helping Robots — and Developers — Move Faster

When Jaiveer Singh talks about robots, he doesn’t begin with spectacle. He begins with infrastructure: the boards inside machines, the software that lets developers see through a robot’s cameras and the engineering required before a robot can leave a demo floor to do something useful. As a robotics software engineer who leads the team behind […]
NVIDIA Blog (AI) 30d ago Agents & autonomy

Shaping AI from the Middle

At New America, our work spans across five issue areas—each essential to building a nation and a world where everyone can thrive. What Parents of Young Kids Want: Insights from the 2026 National ...
New America 30d ago Children & education

Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning

Editor’s note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners, and enterprises can transform their workflows using the latest advances in OpenUSD and NVIDIA Omniverse. Vision AI agents are becoming a practical way to automatically turn video data from the physical world into operational intelligence in factories, […]
NVIDIA Blog (AI) 30d ago Agents & autonomy

Press Regulation: Panic at The Telegraph over fears Burnham will put the public above newspaper owners – Nathan Sparkes

Since the Leveson Report was published in 2012, exposing a collapse in ethical standards across the press, most national newspapers have adopted a similar stance: objection to the very principle of accountability. They believe that while social media, broadcast media and every other industry should be regulated, they alone should be permitted to operate and […]
Inforrm (media law) 30d ago RegulationTransparency

Policy (3)

Research (68)

SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora capture how instructors actually deliver feedback in real classrooms, and no reliable method measures whether generated feedback aligns with what an instructor would write. We address both. SEFORA is a public corpus pairing instructor inline feedback with assignment prom
arXiv 29d ago Jobs & economyChildren & education

Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows

LLMs increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidence that triggered a repair. We introduce Agentic Transaction Processing (ATP), a transaction model that treats generated actions as untrusted proposals until they pass deterministic admission under a declared, executable constraint set C. The governing principle is two-sided: a proposal is not truth, and no proposal for
arXiv 29d ago Agents & autonomy

A Category Theory Account of AI Identity

Artificial intelligence (AI) systems are routinely modified after deployment through retraining and changes in their environments. These transformations raise a metaphysical question: under what conditions does an AI system remain the same system over time or across deployments? Earlier work formulates synchronic and diachronic identity propositionally, by relating identity within a fixed AI system type to equality of trustworthiness levels. Such criteria specify when identity statements are tru
arXiv 30d ago Environment

SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing

Reinforcement learning for diffusion large language models (dLLMs) has largely moved to trajectory-aware methods. The current state of the art, TraceRL, holds that random masking is mismatched with the model's inference trajectory, and it reconstructs that trajectory during training by slicing each rollout into up to K/s trajectory-aligned training samples, a cost that grows with the block size K. We show that this mismatch can be mitigated without reconstructing the trajectory. Our method, SLIM
arXiv 30d ago

A Mechanism-Driven Theory of Phase Transitions in Active Learning

Active learning (AL) performance is known to be budget-dependent, yet regimes are typically defined by heuristic label counts that fail to generalize across datasets or architectures. We characterize AL dynamics by reframing budget regimes as shifts in the dominant generalization mechanism. By reinterpreting PAC-style risk components as dynamic interacting terms, we prove that dominance shifts are structurally unavoidable, creating a moving bottleneck for generalization. We operationalize this u
arXiv 30d ago

Would You Marry Superintelligence?

Emotional bonds between humans and AI companions are growing, and the question of whether a person may marry an AI system will soon move from speculative fiction into law. This chapter examines whether the autonomy-centered logic that has expanded marital choice among human beings can justify extending marital status to superintelligent companions. Following a scenario-envisioning exercise informed by anticipatory ethics, I argue that granting such status leads to socially unjust outcomes, even
arXiv 30d ago RegulationAgents & autonomy

SNAP-FM: Sparse Nonlinear Accelerated Projection for Physics-Constrained Generative Modeling

Generative models have emerged as scalable surrogates for physical simulation, yet they offer no guarantee that their outputs respect the conservation laws, boundary conditions, and nonlinear invariants that govern the underlying physics. Constrained sampling closes this gap, enforcing such constraints exactly at inference time without retraining, but at a computational cost: projection, correction, and trajectory-optimization steps are repeated during sampling, with these steps becoming expensi
arXiv 30d ago

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-only rewards provide too sparse guidance, failing to inform the model about the goodness of intermediate actions. Dense supervision methods aim to solve this problem by scoring intermediate steps, from intrinsic confidence to self-distillation and embedding similarities. However, it is common practice to evaluate them by measuring the downstream perfo
arXiv 30d ago Agents & autonomy

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes. Yet LLMs exhibit systemic deficiencies in key metacognitive faculties: they hallucinate with high confidence, fail to recognize knowledge boundaries, and misrepresent their internal uncertainty--undermining trustworthiness and reliability. Since monitoring task performance and adapting behavior accordingly are central to metacognition, we posit that models capab
arXiv 30d ago Regulation

Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models

Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing the support set for task induction. In the standard ICL setting, once the task is induced, its decision criterion remains fixed. However, in real-world applications, many tasks exhibit a stable high-level intent, while their decision criteria shift according to specific requirements. Thus, we introduce a new setting, denoted as Criterion-Conditional In-Conte
arXiv 30d ago

MVP-Nav: Multi-layer Value Map Planner Navigator

Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment. Existing approaches either rely on high-level semantic reasoning without geometric grounding or learn end-to-end policies that lack explicit physical constraints, often resulting in semantically plausible but physically unsafe behaviors. In this paper, we propose
arXiv 30d ago Safety & alignmentAgents & autonomy

Harnessing Textual Refusal Directions for Multimodal Safety

To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their unimodal counterpart. In this work, we relax this constraint and investigate whether textual refusal directions, extracted directly from the LLM backbone, generalize across modalities (i.e., image, video). Preliminary f
arXiv 30d ago Safety & alignmentFinance, VC & PE

Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling

A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observable closed-loop environments. In such logs, the ego vehicle has rich local observations, while surrounding agents are only partially observed due to perception limits and occlusions. As a result, simulators may learn incomplete context--action mappings that remain hidden in log-based training but emerge during closed-loop rollouts, leading to unreal
arXiv 30d ago Agents & autonomyEnvironment

Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision

Despite recent progress, the reasoning capabilities of large multimodal language models (MLLMs) remain fundamentally constrained by static supervision, where fixed prompts, rules, or reward models provide non-adaptive guidance throughout training. Such static signals are often sufficient to enforce output formats, but fail to shape the underlying reasoning process, leading to brittle generalization and performance saturation in complex decision-making tasks. We propose Evo-PI, a principle-centri
arXiv 30d ago Healthcare

A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical execution. We developed ProtoPilot, a self-evolving multi-agent system, together with an expert-grounded benchmark and evaluation framework for testing this conversion as an experimental automation problem. The framework spans 294 synthetic-biology and molecular
arXiv 30d ago Jobs & economyAgents & autonomy

FedXDS: Leveraging Model Attribution Methods to counteract Data Heterogeneity in Federated Learning

Explainable AI (XAI) methods have demonstrated significant success in recent years at identifying relevant features in input data that drive deep learning model decisions, enhancing interpretability for users. However, the potential of XAI beyond providing model transparency has remained largely unexplored in adjacent machine learning domains. In this paper, we show for the first time how XAI can be utilized in the context of federated learning. Specifically, while federated learning enables col
arXiv 30d ago Safety & alignmentTransparency

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could be shared from what has been shared between dialogue participants through grounding. We formulate this as an interpretation-matching task on 13,077 annotated reference expressions from HCRC MapTask dialogues, and evaluate VLMs under systematically controlled manipulation
arXiv 30d ago Finance, VC & PE

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, on which top-tier systems already achieve near-perfect scores. As T2I models enter creative workflows, users issue multi-faceted requests combining intricate spatial relationships, stylistic constraints, and complex text rendering. In this setting, a single binary V
arXiv 30d ago

When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection

Feature rankings are widely used in supervised feature selection because they are simple, scalable and easy to interpret. Variables are first ranked by a relevance score, and a subset is then obtained by retaining the top-ranked variables. Although the first stage has been extensively studied, the second is often governed by an arbitrary cardinality, an empirical threshold or cross-validation, without a direct interpretation. This raises a basic question: given a feature ranking, when is there e
arXiv 30d ago

Histogram-constrained Image Generation

Diffusion models have emerged as a dominant paradigm in generative modeling, enabling high-fidelity sampling from complex data distributions. Despite impressive capabilities, controlling diffusion models to produce outputs aligned with user intent remains an open challenge, especially when balancing global coherence with local precision. Existing control mechanisms vary in the granularity of their conditioning signals. For example, textual prompts guide generation globally through high-level sem
arXiv 30d ago

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behaviors are genuine or superficial. We show that current fairness evaluations substantially overestimate moral safety. Models appear fair when demographic identity is stated as an explicit label, yet become measurably less fair when the same identity must be inferred. We term this failure performative compliance, where a model is fair when the present
arXiv 30d ago Bias & fairnessRegulation

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, coding environments, robotic systems, security-operation workflows, and autonomous agents that can read private data, call tools, write files, execute code, and act across organizational boundaries. This shift changes the security problem: risks do not arise from the model weights alone, but from the full lifecycle and application stack through which data, promp
arXiv 30d ago Agents & autonomyEnvironment

Scientific Explanations in Health Sciences: Causality, Trust, and Epistemic Adequacy

Medical Artificial Intelligence (AI) is widely expected to transform clinical practice, yet the decision-making processes of many Machine Learning (ML) models remain opaque. Explainability has been advanced as a partial remedy to clarify why AI generates predictions, particularly in high-stakes contexts. Despite ongoing efforts, debates on what constitutes an adequate medical explanation remain unsettled. Yet, explanation has long been a central topic of inquiry in the philosophy of science and
arXiv 30d ago HealthcareTransparency

Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models

Engineering specifications such as interlocks, alarm rationalization tables, and cause-and-effect (C&E) matrices remain central to process control and safety, yet their creation is still predominantly manual, document-driven, and prone to inconsistency. This paper presents a semantic-AI framework that automates the generation of C&E logic by combining a knowledge graph (KG) with a constrained large language model (LLM) layer. The KG builds on an established modular alignment ontology to represen
arXiv 30d ago Safety & alignment

Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation

Radar sensors provide reliable perception under adverse weather and lighting conditions, but their sparse, noisy, and weakly semantic measurements make dense semantic segmentation challenging. Most existing radar segmentation methods rely on grid-based encodings and pairwise interactions, which struggle to capture the higher-order relational structure formed by multiple radar returns from the same physical object. We introduce a unified higher-order structural alignment framework for multi-view
arXiv 30d ago Safety & alignment

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment

Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insecure code, leads to broadly misaligned behaviour on unrelated prompts. Previous work has noted that the severity of EM is highly sensitive to training choices; however, we still lack a systematic characterisation of this sensitivity. We perform a sweep over several Qwen3 models, optimisers, datasets, and batch sizes, and find that the choice of optimiser has t
arXiv 30d ago Safety & alignment

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models

Audio-Language Models (ALMs) achieve strong zero-shot performance by aligning audio with textual class descriptions. Although prompt learning improves accuracy on base classes through few-shot supervised adaptation, we observe a critical trade-off: it often degrades performance on novel classes, sometimes falling below zero-shot accuracy. This exposes a base-to-novel generalization gap in prompt learning for ALMs. To address this issue, we propose \textbf{ZEBRA} (Zero-shot Entropy-Regularized Pr
arXiv 30d ago

FLARE-AI: Flaw Reporting for AI

Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety. Yet the AI reporting ecosystem is fragmented: researchers who identify flaws often do not know what or where to report, and groups who receive reports rarely share them with other relevant stakeholders. As a result, good-faith reporters duplicate effort by submitting many different forms, and recipients lack standardized, triage-ready information. We audit 12 reporting systems published
arXiv 30d ago Safety & alignmentTransparency

A time-series classification framework for individual-level absenteeism prediction under severe class imbalance

Staff absenteeism imposes substantial operational costs in high-demand work environments such as healthcare, emergency services, meat processing, construction, and courier and delivery services, where proactive workforce planning depends on reliable individual-level absence prediction. Existing regression and classification approaches share a structural limitation; they map features observed at time t to labels at the same time t, reproducing already-realised outcomes rather than predicting futu
arXiv 30d ago Jobs & economyHealthcare

On the Convergence of Self-Improving Online LLM Alignment

The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL has demonstrated strong performance on this task. However, a formal analysis of its convergence properties has been lacking. We identify a key theoretical challenge: the standard SAIL objective function is not guaranteed to be strongly concave due to unfavorable properties of its Hessian. To address this limitation, we
arXiv 30d ago Safety & alignment

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment. In practice, however, as market context accumulates over long horizons, these mandates gradually lose their behavioral influence, a phenomenon we formalize as Mandate Salience Decay (MSD). To measure MSD objectively, we introduce FinPersona-Bench, a
arXiv 30d ago Agents & autonomy

Additive Causal Construction for Transferable and Reconfigurable Cross-System Learning in Multi-Source Image Fusion

In multi-source image fusion scenarios, heterogeneous inputs are typically driven by distinct generative mechanisms and can be viewed as a composition of multiple causal systems. However, cross-system discrepancy (CSD) and cross-system entanglement (CSE) commonly arise during the fusion process, often leading to significant performance degradation under out-of-distribution (OOD) predictions. To address the CSD and CSE issues, we propose the additive causal construction (ACC) framework, which cha
arXiv 30d ago

Von Mises Based Uncertainty Quantification for Closely Spaced Automotive Radar Targets

This work investigates uncertainty-aware deep learning approaches for direction of arrival (DOA) estimation in automotive radar, focusing on probabilistic modeling and downstream integration. A circular-statistics-based von Mises (VM) ensemble (ENS) is compared with an evidential deep learning (EDL) framework based on a normal inverse gamma formulation, yielding a Student t predictive distribution in the Euclidean domain. The ENS framework produces angular predictions parameterized by (mu, kappa
arXiv 30d ago Children & educationFinance, VC & PE

CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift

Cloud virtual machines are often overprovisioned, creating avoidable cost and operational inefficiency. We present CLOUDADV, an interactive engineer-facing advisory system for cloud instance sizing under workload drift. The system combines zero-shot time-series forecasting with bounded recommendation generation across day-, week-, and month-scale planning horizons. For each query, CLOUDADV constructs a structured decision context from historical utilization, forecast summaries, current VM metada
arXiv 30d ago

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation

Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities. However, existing research rarely extends this paradigm to the tactile domain, where both object-level semantics and sensor-level configurations jointly determine the meaning of touch. To address this gap, we propose UniTac, the first UMM designed for tactile understanding and generation. UniTac models the tactile process as a transition from non-contact to contact, ca
arXiv 30d ago

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

Artificial intelligence is transforming our capability to solve biological challenges. In dimensionality bottleneck regimes exacerbated by high-dimensional biological data, neural networks force distinct concepts into the lower dimensions known as superposition. Although this superposition is widely known to hinder interpretability, its impact on corrupting the geometry of latent spaces remains critically overlooked. Here, we utilized sparse autoencoders (SAEs) trained on over 100,000 multiplexe
arXiv 30d ago Safety & alignmentHealthcare

Stage-Transition Dense Reward Modeling for Reinforcement Learning

Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations. This work proposes Stage-Transition Dense Reward (STDR), a visual reward-learning framework that converts unstructured expert videos into logically grounded dense rewards for training RL agents from scratch. STDR leverages semantic understanding to infer a task's stag
arXiv 30d ago Agents & autonomyEnvironment

PGUDA: Pressure-Guided Unsupervised Domain Adaptation with Cross-Modal Knowledge Distillation for sEMG-Based Gesture Recognition

Surface electromyography (sEMG)-based gesture recognition has emerged as a promising technology for natural human-computer interaction. However, its practical deployment remains challenging due to severe performance degradation caused by feature distribution discrepancies across different subjects and recording sessions. Although domain adaptation (DA) techniques are commonly employed to mitigate such discrepancies, conventional methods often struggle to effectively aligning sEMG features, prima
arXiv 30d ago

Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models

Joint OFDM-RIS optimization for 6G is a mixed-integer nonlinear programming (MINLP) problem covering sum-rate maximization, energy efficiency, max-min fairness, and peak-to-average power ratio (PAPR)-constrained objectives. Seventy-eight joint OFDM-RIS optimization works published between 2021 and 2026 are surveyed. No standardized benchmark exists, and cross-paper comparisons remain infeasible. This survey classifies these works into four paradigms: (I) model-based convex relaxation, (II) heuri
arXiv 30d ago Bias & fairnessEnvironment

HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)

We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic. Designed in collaboration with a historian, the corpus captures complex reasoning patterns typical of historical inquiry, including cross-source synthesis, temporal reasoning, and the integration of sparse evidence. The dataset is made of 1782 questions and emphasizes multi-hop connections across heterogeneous historical d
arXiv 30d ago

Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling

Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry. We present Embodied CAD, solver-grounded LLM agents for parametric B-Rep assembly modeling. Instead of generating a complete script in one pass, the agent iteratively selects actions from a strati
arXiv 30d ago Agents & autonomy

Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law

Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards focus narrowly on verbatim memorisation. EU copyright doctrine applies a broader standards: substantial similarity, which extends to stylistic choices, narrative structure, and creative elaboration. This mismatch between what current methods detect and what the law protects leaves a significant compliance gap. We introduce PSALM, an LLM-as-a-judge framework that
arXiv 30d ago RegulationCopyright & IP

Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?

As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing values. However, existing research on LLMs with moral dilemmas overlooks a central aspect of human moral cognition: the ability to imagine alternatives that move beyond the given options. We introduce MoralAltDataset, a dataset of 307 moral dilemmas spanning narrative Advisor dilemmas and AI-facing Agent dilemmas, each augmented with compromise and reframed
arXiv 30d ago Agents & autonomy

Long-term Traffic Simulation via Structured Autoregressive Modeling

Interactive traffic simulation is a vital world model for autonomous driving. A central challenge in long-horizon simulation is modeling sustained multi-agent interactions, which is further exacerbated by dynamic token cardinality as agents continuously enter and exit the scene. In this work, we propose that the solution lies in the synergy between the architectural inductive biases and statistical priors of large-scale sequence models, e.g., Large Language Models (LLMs). Our probing experiments
arXiv 30d ago Agents & autonomy

Distilling Temporal Coherence into 2D Networks for Transrectal Ultrasound Prostate Video Segmentation

Real-time video segmentation of the prostate in Transrectal Ultrasound (TRUS) is essential for image-guided interventions. While conventional 2D methods suffer from inter-frame inconsistencies by disregarding temporal context, 3D architectures incur prohibitive latency. To resolve this dilemma, we present a Temporally Consistent Learning Framework that distills temporal coherence into a 2D network during training, preserving single-frame inference efficiency. Our design is driven by a key clinic
arXiv 30d ago

Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation

Adaptive experiments for average treatment effects (ATE) require randomized allocations balancing valid inference with statistical efficiency. The oracle design is a covariate-dependent Neyman rule governed by unknown arm-conditional outcome variances. We investigate whether this sequential variance-estimation and allocation process can be amortized via in-context learning. We introduce Bayesian in-context experimenters: transformer policies trained to imitate a Bayesian posterior Neyman teacher
arXiv 30d ago Finance, VC & PE

AETDICE: Unified Framework and Offline Optimization for Nonlinear Multi-Objective RL

Optimizing nonlinear preferences in multi-objective reinforcement learning (MORL) is essential for capturing complex trade-offs like risk aversion or fairness. However, such non-linearity has historically bifurcated nonlinear MORL objectives into two distinct paradigms: Scalarized Expected Return (SER) and Expected Scalarized Return (ESR). While SER requires global-level optimization and ESR requires non-Markovian policies, leading to fragmented optimization strategies, we bridge this divide thr
arXiv 30d ago Bias & fairness

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps between high-level instructions and low-level motor commands, and inference inefficiency due to autoregressive scalar decoding. In this work, we propose MIRTH, a unified framework designed to address these challenges. MIRTH
arXiv 30d ago Agents & autonomy

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding

3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions. Existing approaches typically perform reasoning over the entire scene, leading to ambiguous predictions and high computational cost, especially in cluttered environments. We observe that many referential expressions rely on local spatial context and often correspond to restricted spatial regions rather than the full scene. Motivated by this insight, we propose PruneGround, an effective pl
arXiv 30d ago Environment

Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records

To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial. Present day simulation based testing methods focus largely on mathematical models for efficient search of optimal scenarios, assuming a fixed scenario representation. On the other hand, real-world testing involves substantial manual effort to design scenario templates for testing. These templates represent distinct failure scenarios consisting of pre-deployment vehicle mo
arXiv 30d ago

Cross-Receiver Open-Set Radio Frequency Fingerprinting via Structure-First Adaptation

Radio frequency fingerprint identification (RFFI) provides a physical-layer credential for Internet of Things devices, but open-set decisions become fragile when a threshold calibrated on a source receiver is applied to a target receiver. Receiver shift can lower the confidence of known transmitters and cause false rejection, whereas closedset alignment can pull unseen target transmitters into known regions and increase false acceptance. This paper presents a Cross-Receiver Open-set Domain Adapt
arXiv 30d ago Safety & alignment

ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs

Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image. In this paper, we identify an internal signature of hallucination: progressive degradation of text-to-image cross-attention during generation, leading to specific failure patterns like unfocused or biased attention. Existing mitigation strategies are largely outcome-driven and do not explicitly target this failure mode. To address this problem, we propose AD
arXiv 30d ago Bias & fairnessSafety & alignment

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this work, we identify a failure mode, deductive stereotyping, in which models apply population-level statistical regularities to individual cases, producing logically coherent yet socially biased inferences. We provide a statistical interpretation of this phenomenon. To steer models toward fairness-aware reasoning, we propo
arXiv 30d ago Bias & fairness

What Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability

Alignment Forum 30d ago Agents & autonomy

NeuroCogMap Reveals Cognitive Organization of Large Language Models

Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal fe
HuggingFace Daily Papers 29d ago Biotech

Investigating LLM-Powered Dissenting Minority Support in Power-Imbalanced Group Decision-Making: Counterargument and Mediation as Intervention Strategies

Minority viewpoints are often suppressed in power-imbalanced group decision-making due to social pressure to comply with the majority. To address this problem, we developed an LLM-powered dissenting minority support system that aimed to foster attention to minority views through either AI-generated counterarguments or AI-mediated messages. We conducted a mixed-method experiment with 96 participants in 24 groups, comparing minority members' experiences across baseline, AI-counterargument, and AI-
arXiv cs.HC 30d ago Finance, VC & PE

The Real Question to Ask About AI Governance

Carolyn Geason-Beissel/MIT SMR | Getty Images Leaders at literally every Fortune 500 company will tell you that they are governing their AI — every single one of them. Now ask those same leaders who’s responsible for shutting down an AI model that’s causing harm. Most people can’t answer that question. That silence is the most […]
MIT Sloan Management Review AI 30d ago Regulation

Reply to Roeser: Nuclear Politics Cannot Ignore Emotions

Roeser (2026) emphasizes the role of emotions in risk-assessment, including risks about nuclear weapons. This is a highly valuable comment. I build upon it with a concrete example: Kenneth Waltz found the risk of nuclear war to be acceptable, but shied away from a world government because of the risk of global civil war. Future philosophical research on nuclear weapons should critically examine the processes behind such judgments about acceptable risk.
Philosophy & Technology 30d ago Military & security

Assertion, Accountability, and Large Language Models

Large language models (LLMs) increasingly participate in communicative practices that resemble human interaction: users ask them questions, rely on their outputs for belief formation and action guidance, and sometimes develop affective attachments. These practices raise a central philosophical question: can the outputs of LLMs be regarded as assertions, and if so, what follows for responsibility and accountability? Standard theories of assertion and testimony assume that assertions require asser
Philosophy & Technology 30d ago Transparency

Artificial Resonance: AI companions as agents of social acceleration

AI companions are becoming increasingly popular, with millions of users worldwide, especially young adults. Some see the potential to fight the so-called loneliness epidemic; others see the destructive effects of addiction and harmful guidance leading users in extreme cases even to suicide. Recent research has examined AI companions through the lens of AI ethics, addressing questions of emotional dependency, controllability and emotional harm. While these contributions are valuable in assessing
AI & Society 30d ago Agents & autonomy

A new paradigm for marine ecological monitoring through swarm intelligence, digital twins, and Human–Swarm interaction

Marine and coastal ecosystems are among the least observable yet most rapidly changing environments, where climate impacts, pollution, and biodiversity loss demand monitoring and intervention at scales that manual sampling and single-robot deployments cannot sustain. This paper argues for a conceptual shift in ecological monitoring and restoration toward networked robotic ecosystems, adopting cooperative swarms of autonomous aquatic robots coupled to in-situ digital twins and human-in-the-loop s
Frontiers in Robotics and AI 30d ago Agents & autonomyEnvironment

Low-cost social robot designs for education: a review

Social robots have shown promising potential in educational contexts worldwide, with studies reporting significant cognitive and affective gains when such robots are deployed. However, among other factors, the high cost of commercial robots limits this line of research to a small number of laboratories and hinders large-scale adoption in real-world educational settings, with most studies remaining short-term pilot interventions. Although several reviews exist in this domain, they primarily focus
Frontiers in Robotics and AI 30d ago Children & educationAgents & autonomy

New AI Flaw Reporting System Fills Crucial Security Gap

Flaw Reporting for AI (FLARE-AI) allows developers and security researchers to submit artificial intelligence flaws for formal, coordinated disclosure.
Carnegie Mellon Software Engineering Institute 30d ago Transparency

Once, cyber-attacks required great skill. AI is changing that.

AI is shrinking the gap between skill and ability for those who want to conduct cyberattacks, BKC Affiliate Bruce Schneier argues in The Guardian. Models merely need user to direction to identify and ...
Harvard Berkman Klein Center 30d ago Military & security

AI and Doctrinal Collapse

Visiting Scholar Alicia Solow-Niederman identifies "inter-regime doctrinal collapse," a source of legal strain by which the boundaries between information privacy law and copyright law become ...
Harvard Berkman Klein Center 30d ago RegulationPrivacy

If an AI chatbot misleads you, who is to blame?

A court in Germany found that Google was responsible for what its chatbots say in search summaries. This is the accountability we need.
Harvard Berkman Klein Center 30d ago Transparency

WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation

The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance disparities across demographic groups. A key obstacle to studying and mitigating such biases is the lack of face detection datasets with sensitive feature annotations. To address this gap, we introduce WIDER-FAIR, a new dataset built on the widely used WIDER-FACE benchmark, manually annotated with the perceived ethnicity and sex of each face. The datase
arXiv fairness query 30d ago Bias & fairness

Comparative Analysis of Machine Learning based Intrusion Detection in Realistic IoT Networks

The Internet of Things (IoT) is rapidly growing and expanding into various sectors, such as healthcare, transportation, smart homes, and more. Despite the benefits of using IoT devices, they present several challenges. Given the significant role these devices play in our lives, it is crucial to address issues related to their security and privacy. These devices are limited in resources, which complicates their security and the protection of the data that they manage. The paper aims to examine in
arXiv cs.CR (AI security) 30d ago PrivacyHealthcare