23:22 UTC
Archive · 2026-07-03

AI ethics on Friday, 3 July 2026

81 items published this day, across 4 categories.

News (11)

AI’s Volatile Power Use Quietly Tests Grid Limits

The rapid expansion of artificial intelligence infrastructure is typically framed as an energy problem. Data centers are projected to consume a growing share of global electricity demand: The International Energy Agency estimates they could account for 3 to 4 percent of total global consumption within this decade. Utilities are already adjusting long-term forecasts to accommodate anticipated growth from hyperscale facilities and high-density compute clusters. This framing captures scale. It miss
IEEE Spectrum 27d ago Environment

The new technologies in the UK defence investment plan

The Defence Investment Plan provides £5 billion for drone warfare.
The Conversation UK Technology 27d ago Military & securityFinance, VC & PE

Parents warned not to publicly share children’s images amid AI abuse risks

The NCA says there is a growing threat of children's images being used to create child abuse material.
BBC Technology 27d ago Children & education

Chinese LLMs Broaden the Gap Between Attackers & Defenders

Two new models from Chinese firms compete with top US mainstream and frontier models. Should cyber-defenders be worried?
Dark Reading (AI security) 27d ago Military & security

“He Didn’t Need to Die.” How an Immigration Detention Center Repeatedly Failed to Address a Mental Health Crisis.

The post “He Didn’t Need to Die.” How an Immigration Detention Center Repeatedly Failed to Address a Mental Health Crisis. appeared first on ProPublica .
ProPublica (Machine Bias) 27d ago Healthcare

The Integrated Circuit and the Future of American AI Leadership

Editor’s note: This is the first article in a limited series celebrating American defense technologies born from wartime and their effects on broader national security, politics, and society. This series will run for several weeks to commemorate America’s 250th anniversary, and winners will be selected by a reader vote undertaken through our newsletter later this summer. Prior installments can be found at the Arsenal of Innovation page. The history of the semiconductor is an origin story for mod
War on the Rocks 27d ago Military & security

How a small Portuguese university city is building hard-tech startups

In the previous episode of Tech Odyssey, we visited Portugal Ventures to see how a public venture capital institution supports early-stage startups in Portugal. This time, we moved further upstream — to Coimbra, a small city where university research, applied R&D, and entrepreneurship are closely connected. Coimbra is better known for history than startups. The […]
TechNode (CN) 27d ago Finance, VC & PE

Misguided and Misunderstood: Trump’s Approach to U.S. Troops in Europe

For many observers, Secretary of Defense Pete Hegseth’s speech on the future of NATO, delivered in Brussels on June 18, 2026, constituted a perfect example of how the Trump administration is angrily abandoning the longstanding U.S. commitment to European security. The prevailing picture is that the administration is eager to shift the burden of Europe’s defense and is thus moving to withdraw U.S. forces from the continent, even though Europe is moving to do more militarily. Hegseth stated, “we’r
War on the Rocks 27d ago Military & security

Kuaishou’s Kling AI raises nearly $3 billion in funding

Kling AI, Kuaishou’s AI video generation business, raised nearly $3 billion in external funding on July 2, with its post-money valuation expected to reach $18 billion. The financing will support Kling AI’s transition to independent commercial operations. The round was co-led by CPE, Guofang Venture Capital, BlueFive, Tencent, CITIC Securities and Zhongguancun Science City Fund […]
TechNode (CN) 27d ago Finance, VC & PE

Tencent joins reported $3 billion funding round for Kuaishou’s Kling AI

Kuaishou’s AI video generation platform, Kling AI, is reportedly close to completing a $3 billion funding round, with Tencent participating in the investment. Upon completion, Kling AI is expected to be valued at $18 billion. In May, foreign media reported that Kuaishou was planning to spin off Kling AI and launch an IPO as early […]
TechNode (CN) 27d ago Finance, VC & PE

Shanghai Boy Defies Fatal Diagnosis to Become a Star Student

A 12-year-old primary school student with a rare muscle-wasting disease gives his peers and teachers a lesson in resilience.
Sixth Tone (CN) 27d ago HealthcareChildren & education

Field notes (7)

Quoting Josh W. Comeau

I just launched my third course, Whimsical Animations, and so far, it’s on track to sell roughly ⅓ as many copies as a typical course launch. It’s a similar story with my two existing courses. Sales are down significantly from last year. There are likely a lot of reasons for this, but I think the biggest is AI. There’s sort of a double whammy with AI: Many people are wondering whether developer jobs will even exist in a few months, so they’re reluctant to spend time/money learning new dev skills
Simon Willisons Weblog 27d ago Jobs & economy

The Fatal Conceit Gets a GPU Cluster: Bernie Sanders’ Plan to Socialize AI

The American A.I. Sovereign Wealth Fund Act rests on a sweeping claim about the ownership of value created by artificial intelligence. Because AI models are trained on data generated by the public, the bill treats the resulting gains as a public resource subject to state control and redistribution. Sen. Bernie Sanders’ (I-Vt.) proposal would require ... The Fatal Conceit Gets a GPU Cluster: Bernie Sanders’ Plan to Socialize AI The post The Fatal Conceit Gets a GPU Cluster: Bernie Sanders’ Plan t
Truth on the Market (digital regulation) 27d ago Regulation

Factories are just rooms

I went into my kid’s school a couple months back and spoke to the year group about manufacturing. Honestly it was the most rewarding speaking gig I’ve done all year. It was about the process of making my AI clock and I have a ton of pics from my factory visit to Shenzhen (mostly pics that I have only shared with Kickstarter backers). I talked about where ideas come from and the value of playing around, and how it’s neat to learn new techniques that you can combine together. I talked about protot
Interconnected (Matt Webb) 27d ago Children & education

How and why to de-Google your life

Plus, a data center rebellion erupts in Canada, Gen Z says it's sexy to be a Luddite, and the push to paint anti-tech activists as extremists. This is Episode 2 of the BITM show with guest Paris Marx.
Blood in the Machine (Brian Merchant) 27d ago Environment

Flock Cameras Can Surveil Cars Without License Plates

This is from a 2024 company presentation : Officers can also tap into data showing a car’s decals, bumper stickers, back and top racks—along with temporary and unique state tags. Flock calls it a “Vehicle Fingerprint” and it’s touted as a way for law enforcement officials to get more information “even when you don’t have full plate information,” the company’s presentation shows. The company gives police officers the ability to search that data as well, to “build stronger cases with less informat
Bruce Schneier — Schneier on Security 27d ago Regulation

Pluralistic: CARDiac, syntax coloring, view source and vibe code (03 Jul 2026)

Today's links CARDiac, syntax coloring, view source and vibe code: With great abstraction comes great power comes great responsibility comes great loss of fidelity. Hey look at this: Delights to delectate. Object permanence: Real elections v reality TV; Copyright troll loses license; Who gets fed housing subsidies? Trump x forced labor. Upcoming appearances: London, Edinburgh, Sydney, Melbourne, Brighton, London, South Bend. Recent appearances: Where I've been. Latest books: You keep readin' em,
Pluralistic (Cory Doctorow) 27d ago Jobs & economyCopyright & IP

Vercel's Andrew Qu on why agents are a new kind of software

The Vercel Chief of Software explains how its agent framework, eve, was created — and why skills, sandboxes and agent-readable websites now matter.
Latent Space 27d ago Agents & autonomy

Policy (6)

European Parliament Plenary Session – July 2026

Parliament's final plenary session before the summer recess will see Members discuss the priorities of the Irish EU Council Presidency alongside key decisions on enlargement, foreign affairs, passenger rights, agriculture, competitiveness, environmental crime and the EU budget.
European Parliament Think Tank 27d ago Environment

Online Safety in Education

This question is for testing whether you are a human visitor and to prevent automated spam submission. What code is in the image? Your support ID is: 4912247370796614757 .
UNESCO AI 27d ago Children & education

Inclusion and anti-discrimination programmes

These activities are directly based on the case law of the European Court of Human Rights, the recommendations and findings of the European Commission against Racism and Intolerance (ECRI), the ...
Council of Europe AI 27d ago Bias & fairnessRegulation

The Health and Economic Benefits of Tackling Non‑Communicable Diseases

Analysis and insights for driving a rapid transition to net-zero while building resilience to physical climate impacts ...
OECD 27d ago HealthcareEnvironment

OECD Digital Education Outlook 2026

The OECD Digital Education Outlook 2026 explores emerging research on the use of generative AI in education and presents innovative tools and applications that show promise. The report examines the ...
OECD 27d ago Children & education

Integrating AI in TVET: A practical guide for institutions

This is creating a gap between how people already use AI and the formal training required to do so ethically and effectively. Technical and vocational education and training (TVET) institutions need ...
UNESCO AI 27d ago Children & education

Research (57)

RADIO1D: Elastic Representations for Condensed Vision Modeling

This paper challenges the assumption that vision-language models (VLMs) require fixed patch-based 2D vision features. Analyzing fine-tuned vision encoders, we find that representations become increasingly abstract and less spatially coherent during VLM training. Notably, models trained with image-text alignment (such as SigLIP2) develop a small number of specialized tokens that effectively summarize global image content. Building on this, we introduce RADIO1D, which compresses images into a comp
arXiv 27d ago Safety & alignment

MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms

Foundation models have recently emerged as a powerful paradigm for learning transferable representations from large scale biomedical data, yet existing approaches for physiological waveforms primarily optimize reconstruction or forecasting objectives that do not explicitly preserve clinically meaningful waveform morphology. Electrocardiograms (ECGs) and pulse oximetry (SpO2) waveforms encode rich cardiovascular and hemodynamic information through their morphological structure. In this work, we i
arXiv 27d ago

Development of a Bio-Inspired Routing Algorithm According to Values of Solidarity and a Freirean Perspective of Engineering

A routing algorithm for Señoritas Courier, a bicycle delivery cooperative in São Paulo, Brazil, composed exclusively of cis women and trans people, is presented in this paper. Unlike conventional logistics optimization, which typically focuses on cost or distance minimization, this cooperative operates under principles of solidarity, care, and equitable income distribution. The algorithm was developed through a participatory process involving cooperative members as co-designers. The classical Ve
arXiv 27d ago

Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems

In distributed systems, the classical State Machine Replication (SMR) model assumes that correct replicas execute deterministic transitions to yield identical bitwise states. However, the rise of agentic distributed systems -- where autonomous, stochastic, and model-driven agents orchestrate infrastructure -- presents scenarios where deterministic, bitwise replication is insufficient. Replicas operating with generative models may exhibit divergent reasoning paths, summaries, and token boundaries
arXiv 27d ago Agents & autonomy

Teacher Supervision over Representation Equivalence Classes

Knowledge distillation is usually framed as a choice of what to match in the teacher - its logits, hidden features, or sample relations - which presupposes that the teacher's representation has absolute coordinates to match. It does not: a pretrained representation is identifiable only up to an orthogonal-and-isotropic-scaling equivalence class, so a student should learn the teacher's equivalence class, not its features. The organizing fact is that capability is the teacher's output function, a
arXiv 27d ago Children & education

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs

As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions. A recent line of work focuses on verification via debate, a model of interactive proofs where two competing powerful provers, or AI models, debate each other to convince a weak verifier, or a human, of the correctness of their claim. However, debate assumes that the two AI models possess equal abilities and that one of them is truthful, which ma
arXiv 27d ago Safety & alignment

Latent Clarity: Bridging World-Model Kinematics to Semantic Manifolds for Video Anomaly Anticipation

Continuous video anomaly detection is dominated by reactive Multiple Instance Learning (MIL) that collapses spatiotemporal features into scalar scores. We introduce PULS (Predictive Unified Latent Space), a continuous semantic world-model pipeline comprising two modules: a 490M-parameter KSD Bridge (Kinematic-to-Semantic Distillation) and a 16.8M-parameter Anticipatory State Predictor (ASP). The KSD Bridge maps V-JEPA 2 physical tensors into the 2048-d Qwen3-VL-Embedding-2B text-aligned hypersph
arXiv 27d ago

PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models

Integrating external tools with Large Language Models (LLMs) has emerged as a promising paradigm for accomplishing complex tasks. Since LLMs still struggle to effectively manage large tool collections, researchers have begun exploring retrieval-based methods to pre-select the most relevant options, addressing input length and latency constraints. However, existing retrievers are often misaligned with tool-calling LLMs due to their separate training processes. This paper presents PORTS, a novel o
arXiv 27d ago Safety & alignment

Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI

Frontier-AI governance today faces a problem structurally analogous to the one banking regulation faced pre-2008, and which post-2008 reforms (Basel III, Dodd-Frank) have since addressed. Two gaps recur: discovering a risk is not tantamount to acting on it, and individual-model review is unlike managing correlated build-up across the sector. Drawing on the Basel III framework and the U.S. financial-stability architecture, I propose a macro-prudential early warning and response system ("MEWRS") f
arXiv 27d ago RegulationFinance, VC & PE

MentalThink: Shaping Thoughts in Mental SVG World

We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink is a think-with-SVG pipeline, where the model learns to generate, render, and interpret scalable vector graphics (SVG) code as an intermediate visual representation for multi-turn reasoning. By creating structured vector sketches, the model can externalize spatial hypotheses, inspect them through deterministic renderin
arXiv 27d ago

Aligning Language Models with Selective Prediction

Large language models (LLMs) are increasingly deployed as critical decision-making components in high-stakes real-world AI systems, rendering LLM reliability a foremost practical concern. In this paper, we focus on enhancing LLM reliability through selective prediction (SP), a strategy that allows an LLM to only predict for inputs where it is likely to be correct (i.e., coverage) and hence reduce the error rate (i.e., risk) on that portion of inputs -- flagging the remaining inputs for future hu
arXiv 27d ago

AGL-1: The Enterprise AI Governance Layer as a Control Plane for Trusted Enterprise Intelligence

Enterprise artificial intelligence is moving from isolated experimentation toward operational dependency across copilots, retrieval-augmented generation systems, autonomous agents, and AI-enabled business workflows. As this transition accelerates, the primary enterprise challenge is no longer only model access or inference scale. It is governed intelligence operations: the ability to enforce authorization, preserve contextual lineage, control persistent memory, detect stale or conflicting knowle
arXiv 27d ago RegulationAgents & autonomy

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI

Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and retrieval-augmented generation, but enterprises are now beginning to deploy agents that plan, retrieve, remember, call tools, update systems, and coordinate work across applications. This changes the evaluation problem. Leaders are no longer asking only whether an answer is accurate or fluent. They need to know who authorized an action, which policy applied, wh
arXiv 27d ago RegulationAgents & autonomy

Demonstrating Generalization Failures via Mixtures of Conditional Policies

Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training and deployment environments. This exposes developers to generalization failures, which are relatively poorly understood. To better understand such generalization failures, we believe the community should construct clean demonstrations under simplified conditions. To facilitate this, we propose a simple and flexible way to construct language models which fail
arXiv 27d ago Environment

Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning

Inference-time alignment methods, such as Best-of-$N$, offer a flexible alternative to training-based alignment by using reward models to select high-quality responses generated by a reference LLM. However, the efficacy of these methods is inherently limited by the response quality: if the reference LLM assigns negligible probability to high-reward responses, no selection strategy will succeed in finding aligned outputs. In this work, we propose Best-of-Better-$N$ (BoBN), an in context learning-
arXiv 27d ago Safety & alignment

AI Systems as Digital Public Goods -- Evidence and Recommendations from a Multi-Stakeholder Assessment

AI systems are increasingly being positioned as potential Digital Public Goods (DPGs) to accelerate progress towards the Sustainable Development Goals (SDGs). Yet, despite major global commitments, most notably the Global Digital Compact's call to "develop, disseminate and maintain safe and secure open-source software, open data, open artificial intelligence models and open standards that benefit society as a whole", very few AI systems currently meet the DPG Standard in practice. This report ex
arXiv 27d ago

Personalized Causal Recourse: A Human-In-The-Loop Approach

Algorithmic recourse addresses the challenge of providing tailored recommendations to users affected by unfavorable machine learning decisions, in potentially high-stakes scenarios. Traditional approaches to recourse often rely on the closest counterfactual explanations or assume a priori knowledge of a user's causal structure, resulting in interventions that overlook individual contexts and specific feature interactions. To overcome these limitations, we study a human-in-the-loop framework that
arXiv 27d ago

Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies

Modern AI agent implementations such as frontier coding agents chain multiple tools at runtime that create a security surface that per-tool guardrails are unable to address, as individually permitted tools can violate organizational policies when composed. We propose the Dynamic Security Control Compositor (DSCC), a two-phase approach to compositional security for multi-tool agent chains. In Phase 1, at session checkout, a Most Restrictive Set (MRS) algorithm composes per-tool security policies
arXiv 27d ago Agents & autonomy

Efficient bias mitigation in T2I diffusion models using Concept Graphs

Text-to-Image diffusion models often propagate harmful bias inherited from the training data. Existing bias mitigation techniques typically intervene only at the text encoder or provide inference-time guidance, often leading to generations that collapse into semantically incoherent outputs. To address these limitations, we introduce CO-ALIGN (Concept Ontology Alignment), a novel bias mitigation approach based on concept-graph alignment that operates on the model's internal concept ontology. By a
arXiv 27d ago Bias & fairnessSafety & alignment

When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions

Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable. We study this problem in a hotel-pricing simulator where an agentic policy editor receives only region-level diagnostic feedback: summaries of how its price distribution differs from a benchmark policy across time, inventory, and market regions. The editor cannot observe benchmark actions, benchmark source code, rewar
arXiv 27d ago RegulationSafety & alignment

Efficient Decentralized Multi-task Dataset Valuation via Model Merging

Accurate and efficient dataset valuation is essential for enabling fair and transparent data marketplaces, especially when multiple contributors provide data for training multi-task models. Most existing valuation methods, however, are limited to single-task settings, overlooking scenarios where a buyer aims to optimize performance across multiple downstream tasks. Moreover, traditional valuation approaches, such as Shapley-based or retraining-based methods, are computationally expensive and poo
arXiv 27d ago TransparencyFinance, VC & PE

A harmonised dataset for Earth system foundation models

Foundation models for Earth systems have so far been trained primarily on physical climate and weather data, with limited representation of the human systems that both drive and respond to environmental change. The lack of a unified global training resource that combines climate, land, ocean, cryosphere, infrastructure, hazards, and socioeconomic data on a common grid hinders progress toward truly multimodal Earth system foundation models. We present WorldTensor, a harmonised global dataset that
arXiv 27d ago Environment

Unbiased Alignment for Large Language Models with Noisy Preferences

The alignment of large language models with human preferences is commonly achieved through Reinforcement Learning from Human Feedback or Direct Preference Optimization. However, these methods are vulnerable to the significant noise prevalent in real-world preference datasets. To address this critical issue, we present a theoretical framework for unbiased alignment, introducing the Unbiased Reward Model (URM) loss and the Unbiased Direct Preference Optimization (UDPO) loss. By mathematically corr
arXiv 27d ago Safety & alignment

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation. Large language models (LLMs) and agentic AI systems, capable of tool use, multi-step reasoning, and iterative intelligence generation, have emerged as promising solutions, yet evaluation frameworks have not kept pace with reported capabilities. This survey systematically reviews 74 studies and makes
arXiv 27d ago Military & securityAgents & autonomy

Organizational Memory for Agentic Business Process Execution

LLM-based agents offer new opportunities for automating business process execution beyond the limits of rule-based systems. However, general-purpose LLMs lack the organization-specific knowledge required for reliable execution, which is typically fragmented across human-oriented artifacts such as policies, process models, and standard operating procedures. While such knowledge can technically be encoded in individual prompts or agent-specific retrieval setups, this approach does not scale in ent
arXiv 27d ago Agents & autonomy

A Bayesian Framework for Evaluating Scenario Compatibility in Generative Population Synthesis

Scenario-based transportation analysis specifies future assumptions through aggregate population targets, whereas generative population synthesis models produce detailed individual-level realizations. When scenario targets are imposed on generative models, current practice relies on deterministic marginal calibration, implicitly assuming that the targets are compatible with the model's learned structural support. However, whether scenario-level constraints lie within the generative support--and
arXiv 27d ago

AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning

Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into executable continuous trajectories. Vision-Language-Action models have been introduced into driving planning to improve long-tail generalization, commonsense reasoning, high-level semantic understanding, and explainability. However, existing VLA planners mainly follow planning-head-based trajectory prediction or full-trajectory autoregressive generation. The for
arXiv 27d ago Transparency

Teaming Up with AI: Coordination and Cooperation

Successful diffusion of AI in the workforce hinges on the economic value that AI brings to human endeavors. Bringing AI into the workforce is more than deploying a powerful new technology -- it is launching a new form of collaboration. Each human worker is now endowed with a team of AI agents; work can be delegated to these agents, and the role of the human shifts towards managing and monitoring. How can we maximize the economic value from collaboration with AI in the workforce? How can we make
arXiv 27d ago Jobs & economyAgents & autonomy

Decentralised Federated Learning over Temporal Networks: The Role of Heterogeneities

Decentralised federated learning, based on peer-to-peer communication, is increasingly proposed for on-device training of machine learning models, promising a privacy-preserving, communication-efficient training process with no risk of single-point failure. However, the role of structural and temporal inhomogeneities in such fully decentralised settings remains poorly understood. Here, we investigate their effects when model parameters are locally averaged during aggregation. We show that the de
arXiv 27d ago PrivacyFinance, VC & PE

KARMA: Knowledge graph-based Automated Reasoning Materialization and Alignment

Template-based contrastive synthesis is scalable, but its candidates often differ only in a few entity-slots while sequence-level optimization spreads supervision over mostly shared templates. We formalize this as the Resolution Mismatch Problem and propose KARMA, which enumerates schema-constrained paths over domain knowledge graphs and verbalizes them into slot-aligned contrastive candidates. Slot-Parallel Alignment (SPA) then applies a decoupled slot-level objective to route preference superv
arXiv 27d ago Safety & alignment

Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning

Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details. We aim to learn representations whose matching is stable across caption views and whose confidence reflects how strongly text constrains an image. We propose Text as Partial Constraint (TPC), a core-residual alignment framework that treats multi-view captions as incomplete s
arXiv 27d ago Safety & alignment

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy

Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult. Existing scalable RL methods typically assign trajectory-level rewards uniformly across tokens, while recent entropy-aware approaches either rely on coarse detached heuristics or directly optimize true entropy, which can introduce non-local gradient components misaligned with sampled-token policy updates. We p
arXiv 27d ago RegulationSafety & alignment

Silicon Sampling via Cross-Survey Transfer

Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and
arXiv 27d ago

Spectral Rewiring for Exploration, Purification, and Model Merging

Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-relevant bottlenecks: suppressed reasoning performance, often reflected by premature saturation of test-time scaling, and interference when consolidating multiple capabilities through multi-domain training or model merging. We show that the reasoning-effective component of these updates is largely concentrated in the base model's spectral space, moti
arXiv 27d ago

OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models

Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate large token sequences under audio-visual inputs, leading to substantial inference cost. Existing audio-visual token compression methods often rely on unimodal guidance, overlooking the temporal locality of query-relevant evidence in audio-visual inputs and implicitly assuming that the two modalities share a temporally aligned information density distri
arXiv 27d ago

Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making

The use of Large Language Models (LLMs) across diverse areas of human activity-ranging from everyday tasks to safety-critical applications-aims to enhance decision-making effectiveness with minimal human feedback. Concurrently, it seeks to align decisions with human expectations, preferences, and needs while mitigating risks associated with AI non-determinism. However, humans frequently over- or under-rely on AI recommendations, and current AI systems remain poorly calibrated to human expectatio
arXiv 27d ago

HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion

Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation. However, their ability to produce long-duration videos is fundamentally constrained by the quadratic complexity of the self-attention mechanism. Recent clustering-based sparse attention methods improve the quality-speed trade-off by grouping semantically similar tokens, but their practical efficiency remains limited by two bottlenecks: substantial clustering overhead and low CTA uti
arXiv 27d ago

Back to Basics: Improving Molecular Understanding in LLMs via SMILES-Graph Translation

Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains often come without reliable structural grounding. In particular, existing approaches conflict with the chemistry principle that structure determines function: despite their downstream success, current molecular LLMs perform poorly on basic structure recognition, suggesting that they fail to capture molecular graphs from canonical SMILES. To remedy thi
arXiv 27d ago

CAFÉ, an automated feedback tool to approach Formal Methods

We present CAFÉ, a learning platform designed to introduce computer science students to Formal Methods (FM). CAFÉ aims to scaffold students' structural thinking (in contrast with operational thinking) by promoting the practice of Graphical Loop Invariant Based Programming (GLIBP). In the GLIBP approach, students solve loop-based problems by first constructing a Graphical Loop Invariant (GLI) before deriving the corresponding code. The GLI is an informal diagrammatic representation of the loop in
arXiv 27d ago Children & education

A Scalable Approach to Evaluating Moral Sensitivity in LLMs

Moral sensitivity is the ability to identify the morally relevant features of a decision situation and use them as the basis for action. It is the foundation of broader moral competence: any other moral reasoning capabilities will be irrelevant if an agent lacks sensitivity to the relevant facts. In this paper, we offer a new evaluation of LLM moral sensitivity and in doing so, we address and resolve a central problem in AI alignment research: how to scale behavioural evaluations beyond expensiv
arXiv 27d ago Safety & alignmentAgents & autonomy

The Foreign Policy AI Evaluation Gap

We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for technical AI governance research. In enacting foreign policy, we refer to the formulation and implementation of external objectives by political actors. Statecraft is a high-consequence deployment domain, with extreme downside risks and structural properties that standard evaluation practices handle poorly. These features include partial observability, unbounded a
arXiv 27d ago Regulation

MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding

Materials phase diagrams are a core knowledge representation in materials science, encoding temperature,composition, phase stability, and phase transformation pathways, with their full understanding requiring thermodynamic mechanism analysis and scientific reasoning. Although VLMs have shown promise in scientific image understanding, their systematic evaluation on such logically complex images demanding deep mechanistic interpretation remains limited, and phase diagrams provide a challenging tes
arXiv 27d ago

Bootstrap Flow-Map Tree Sampling Enables Online Feedback Driven Search

In many scientific and engineering domains, maximizing discovery within a limited sampling budget demands strategic, observation-guided exploration. While generative models have enabled training-free reward alignment, current methods typically excel in local searches within narrow regions of the underlying distribution. These approaches struggle when preferences are unknown a priori and only revealed through sequential feedback-a scenario demanding broad exploration to uncover high-utility regio
arXiv 27d ago Safety & alignment

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains a persistent challenge. Conventional refusal-oriented alignment strategies mitigate harmful content generation but systematically fail to serve legitimate user needs, often withholding information that could safely and constructively address the underlying intent of sensitive queries. Building upon the constructive sa
arXiv 27d ago Safety & alignment

TIER: Trajectory-Invariant Explanation Regularization for Membership Privacy

Explainability is central to building trustworthy AI, yet explanation interfaces can inadvertently provide adversaries with an expanded privacy-related attack surfaces. Recent studies show that advanced membership-inference attacks succeed by exploiting confidence-drop trajectories, induced through attribution-guided perturbations, as discriminative features, rather than directly using confidence scores or explanation vectors. Existing defenses against membership inference fail to directly mitig
arXiv 27d ago Bias & fairnessPrivacy

SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal Roughness

Deploying AI-generated video detectors in real-world services demands an ultra-low false positive rate (FPR) on real videos to avoid falsely rejecting authentic content, a regime where standard metrics such as AUROC fail to reflect actual operating behavior. We introduce Spatial Patch-Level Incoherence and Temporal Roughness (SPLIT), a training-free detector that operates on patch tokens from a frozen vision encoder to detect both fully generated and partially edited videos. SPLIT computes two c
arXiv 27d ago

Pragmatic FDT, and predictors as game theory

Alignment Forum 27d ago

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

A key challenge in Arabic NLP is the scarcity of dialectal data relative to Modern Standard Arabic (MSA), causing LLMs to overproduce MSA and struggle with dialectally accurate generation. From an interpretability perspective, this raises a fundamental question: where and how are dialectal features encoded within model internals, and can these representations be leveraged to improve dialect generation without fine-tuning? This study investigates two complementary inference-time approaches that s
HuggingFace Daily Papers 27d ago Safety & alignmentFinance, VC & PE

Regulating AI: Where U.S. State Policy and HCI (Mis)align

Artificial intelligence (AI) technologies are increasingly adopted into everyday life, with most investment and development concentrated in the U.S. In response to rapid AI integration and scant federal guidelines, U.S. states have formed AI committees charged with studying AI-related societal trade-offs. We analyzed the 18 existing state-level AI committee reports to understand how policymakers discuss AI-related benefits and risks. We then compared the risks surfaced by policymakers to an esta
arXiv cs.HC 27d ago RegulationFinance, VC & PE

Using AI, Using Humanity: A Response to Sticker

Martin Sticker (2026) expresses sympathy for our application of Kant’s formula of humanity to the morality of LLMs (Aylsworth and Castro 2024), but he also raises objections: he claims that the argument rests on an ambiguous conception of “humanity;” that it entails the absurd conclusion that every student ought to specialize in the humanities; and that we should think of LLM use in terms of the prohibition to use others as a mere means (rather than in terms of a self-regarding duty to cultivate
Philosophy & Technology 27d ago Children & education

A review on empirical studies in explainable artificial intelligence

As artificial intelligence (AI) systems become more integrated into decision-making processes, the need for explainability has emerged to foster trust, understanding, and effective human-AI collaboration. With the variety of explainable AI (XAI) methods available, selecting the right one for a specific user group and a given use case remains challenging, especially given the limited empirical validation of existing theoretical guidance. This systematic literature review addresses this gap by syn
Artificial Intelligence Review 27d ago Transparency

Business Ethics and Ethical Business

This second edition of Business Ethics and Ethical Business has all the virtues of the first edition as well as additions and revisions throughout, and three new chapters—one on AI and two on environmental ethics and sustainability. It explores the place of business in society, describes ethics in management, recognizes the challenges of environmental preservation, and explores challenges of international business. It introduces the major standards of business ethics and provides tools for ethic
OpenAlex 27d ago Environment

CodePori: Large-Scale System for Autonomous Software Development Using Multi-Agent Technology

OpenAlex 27d ago Agents & autonomy

Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

Detecting mental health disorders in a timely manner is an important societal challenge. NLP and machine learning (ML) methods used to assist with detection rely on data collected primarily from social media. However, such datasets often have sampling biases and inherent ethical and privacy issues. One avenue to overcome these limitations is non-social media data. We present the first comprehensive review of non-social media, free-text datasets for mental health research. We use the PRISMA metho
arXiv cs.CL (ethics-relevant NLP) 27d ago PrivacyHealthcare

CONTRA: Red-Teaming Configurations of Personalizable Agents

Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These systems allow personalization of the agent through modifiable internal files and the installation of skills. While this enables deployment in a wide range of settings and the automation of diverse tasks, greater capability and autonomy increases the risk of malicious actions being executed unintentionally. In this work, we explore the interplay betwe
arXiv red teaming query 27d ago Safety & alignmentJobs & economy

Builder, Defender, Breaker: The Case Against Removing the Human from the AI-Driven Security Lifecycle

Artificial intelligence has spread across the whole of the security lifecycle. The same family of models now writes application code, hardens it, and probes it for weaknesses, so that a single generative substrate increasingly performs all three roles at once. Enthusiasm for this convergence tends to treat full autonomy as the natural end point of partial assistance. This article argues that it is not. When the system that builds an artifact is drawn from the same distribution as the systems tha
arXiv cs.CR (AI security) 27d ago Agents & autonomy

Overloading Large Vision-Language Models for Jailbreaking

Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as personal assistants, document analysis systems, and embodied agents. However, their dual-modal attack surfaces make them vulnerable to jailbreak attacks. Existing LVLM jailbreaks rely on simple designs, e.g., short text and out-of-distribution images. Nevertheless, recent advancements in both large language model backbones and multimodal mechanisms
arXiv red teaming query 27d ago Safety & alignmentAgents & autonomy