23:23 UTC
Archive · 2026-06-25

AI ethics on Thursday, 25 June 2026

87 items published this day, across 5 categories.

Incidents (4)

NewsBreak: Most downloaded US news app has Chinese roots and 'writes fiction' using AI

LONDON, June 5 (Reuters) - Last Christmas Eve, NewsBreak, opens new tab, a free app with roots in China that is the most downloaded news app in the United States, published an alarming piece about a small town shooting. It was headlined "Ch ... (https://incidentdatabase.ai/cite/1554#7451)
AI Incident Database 35d ago

Citation errors and hallucinated case turn up in Boies Schiller brief in 'artificial-intelligence debacle'

A partner at Boies Schiller Flexner is seeking to file a corrected brief in litigation against the Church of Scientology after taking responsibility for "material citation errors" that it contained. In a Sept. 19 declaration, partner John ... (https://incidentdatabase.ai/cite/1555#7453)
AI Incident Database 35d ago

X user tricks Grok into sending them $200,000 in crypto using morse code

An X user managed to trick AI chatbot Grok into sending around $200,000 worth of crypto after exploiting its link with an automated trading bot. The incident involved Grok and 'Bankrbot', two AI systems with wallet access, which were manip ... (https://incidentdatabase.ai/cite/1556#7454)
AI Incident Database 35d ago

Former City Council Candidate Charged With Forgery for Disseminating Altered Political Endorsements and Phony News Reports

Queens District Attorney Melinda Katz announced that Jonathan Rinaldi has been charged with forgery and criminal possession of a forged instrument for allegedly creating and distributing false political endorsements and fake news articles u ... (https://incidentdatabase.ai/cite/1557#7455)
AI Incident Database 35d ago Misinformation

News (13)

For Students, the Process of 'Becoming' is the Challenge No Chatbot Can Solve

Tech Policy Press 35d ago Children & education

Who Really Controls Your Digital Likeness in the Age of AI Wearables? Not You.

Tech Policy Press 35d ago

Google and Apple’s Anti-DMA Lobbying Strategy Goes All-in on Security and Privacy

Tech Policy Press 35d ago Privacy

They Said I'd Feel Different About Free Speech as a Parent. They Were Wrong.

Tech Policy Press 35d ago

How Will Andy Burnham Handle Tech Policy as UK Prime Minister?

Tech Policy Press 35d ago Regulation

India’s Telegram Ban Was Temporary. The Power Behind It Is Not.

Tech Policy Press 35d ago

Hugging Face hosts nudification tools targeting a former Trump cabinet official and other senior US political figures

The tools are explicitly intended for generating deepfake nudes of a former Trump cabinet official, sitting members of Congress and a top American judge, a Transformer investigation found
Transformer 35d ago MisinformationFinance, VC & PE

Improving the speed and energy-efficiency of AI agents

A new system, known as Murakkab, optimizes the design and deployment of multistep workflows that power AI applications.
MIT News 35d ago Agents & autonomyEnvironment

STAT+: At BIO 2026, industry wrestled with Washington politics, and making AI work better

Biotech executives reveal concerns over Chinese biotech, the profitability of AI, and the durability of Trump's drug price moves.
STAT News (health AI, headlines) 35d ago HealthcareBiotech

The MAHA Movement’s Worrisome Embrace of Ibogaine

The MAHA movement has been promoting the plant-based substance ibogaine as a remedy for opioid addiction, despite the lack of clinical trials on its safety and efficacy. One physician who practices addiction medicine discusses the risks of ibogaine and argues for the need to support proven treatments.
Undark Magazine 35d ago Healthcare

AI backers wanted a knockout win in New York. Now they’re clamming up.

Leading the Future spent big to stop the author of the state’s AI safety law from going to Congress. Then came the backlash.
Politico Technology (US) 35d ago RegulationSafety & alignment

$500 million AI jobs push launches with bipartisan backing

A new bipartisan group will work with corporate donors like Anthropic, OpenAI, Amazon, Microsoft and Bank of America to retrain workers displaced by the AI boom.
Politico Technology (US) 35d ago Jobs & economy

IEC dreams of digital voting, just not in this election

The IEC is gearing up to deliver a more technology-enabled process at the polls later this year.
ITWeb (ZA) 35d ago Misinformation

Field notes (13)

AI Bias Is Putting LGBTQIA+ People at Risk

The post AI Bias Is Putting LGBTQIA+ People at Risk appeared first on Partnership on AI .
Partnership on AI 35d ago Bias & fairness

How agents are transforming work

A new OpenAI research paper shows how AI agents are transforming work, enabling longer, more complex tasks and expanding productivity across roles.
OpenAI 35d ago Jobs & economyAgents & autonomy

EFF, TEDIC and CEJIL Challenge Secrecy in the Use of Face Recognition in Paraguay

Seeking transparency and accountability in Paraguay’s use of facial recognition, EFF, the Association of Technology, Education, Development, Research, Communication (TEDIC), and the Centre for Justice and International Law (CEJIL) filed a complaint with the Inter-American Commission on Human Rights against the state for arbitrarily denying access to information about its implementation and use of the technology as a tool for mass surveillance that erodes people’s privacy rights. The case involve
EFF Deeplinks 35d ago RegulationPrivacy

Four Years After Dobbs, Anti-Abortion Lawmakers Keep Coming for Online Speech

This week marks four years since Dobbs v. Jackson Women’s Health Organization overturned Roe v. Wade ’s constitutional protections for people seeking abortion care. Anniversaries are a moment to take stock, and over the last four years, EFF has seen firsthand how digital rights and reproductive rights have become increasingly intertwined. One major way this has happened: the fight over abortion has also become a fight over online speech and government censorship as a steady stream of proposed la
EFF Deeplinks 35d ago Healthcare

The FCC’s Spam Call Proposal Is Just a Data Collection Scheme

The Federal Communications Commission wants to require telecommunications providers to collect vast amounts of personal information from every person who wants a phone number in the name of combatting scam and spam calls. This plan will fail to combat the deluge of unwanted calls people in the United States receive every day while giving untrustworthy companies a gold mine of information that would harm everyday consumer’s privacy, access to communications, and ability to speak freely. The requi
EFF Deeplinks 35d ago Privacy

Are Your Local Police Using Flock Safety ALPRs to Scan for Immigrants?

When a car passes an automated license plate reader (ALPR), its plate is captured and instantly compared against a list of vehicles that police are actively looking for or that police have identified for real-time surveillance. These are called “hotlists,” and EFF has learned that one used by agencies across the country targets immigrants on behalf of Immigration and Customs Enforcement (ICE). Agencies using Flock Safety ALPR systems commonly allow the plates their cameras collect to be compared
EFF Deeplinks 35d ago Privacy

Russia used Cellebrite tool to jail activist after company claimed to have ended contract

A Citizen Lab investigation confirms evidence Russian authorities used Cellebrite tool to hack prominent Russian activist Andrey Pivovarov after the Cellebrite claimed to have ceased sales to Russia. The post Russia used Cellebrite tool to jail activist after company claimed to have ended contract appeared first on Access Now .
Access Now 35d ago Finance, VC & PE

The KIDS Act Would Require Age Checks To Get Online

Within the next week, Congress is preparing to vote on the KIDS Act , a sprawling package of legislation that seeks to control Americans’ web browsing and private messaging. The package includes a revised version of the Kids Online Safety Act , or KOSA, combined with a collection of other internet bills, study bills, reporting requirements, and new regulations. Instead of debating any of these proposals on their merits, lawmakers are attempting to move them all at once under an ultra-expedited p
EFF Deeplinks 35d ago RegulationChildren & education

Global Freedom of Expression, Columbia University: Newsletter, 25 June 2026

Columbia Global Freedom of Expression seeks to contribute to the development of an integrated and progressive jurisprudence and understanding on freedom of expression and information around the world. It maintains an extensive database of international case law. This is its newsletter dealing with recent developments in the field. “I am certain that the machinery of violence […]
Inforrm (media law) 35d ago Regulation

The Anti-SLAPP Bill: an unfocused invitation to expense and abuse – Hugh Tomlinson KC

On 16 June 2026 Baroness Stowell introduced the Strategic Litigation Against Public Participation Bill (“the Bill”) into the House of Lords. An identical bill has been introduced in the House of Commons by Sir John Whittingdale MP. Unfortunately, despite its title, the Bill is not focussed on the issue of “SLAPPs” or abusive litigation. Rather […]
Inforrm (media law) 35d ago Regulation

🔮 The state of the AI economy

We've reconstructed the AI economy from the bottom up
Exponential View (Azeem Azhar) 35d ago Jobs & economy

Pluralistic: Jailbreaking isn't theft (25 Jun 2026)

Today's links Jailbreaking isn't theft: It wasn't progress when they did it, it's not piracy when we do it back to them. Hey look at this: Delights to delectate. Object permanence: Major AI breakthrough; Disney v Pooh tombstone; Vancouver riot kiss; Farage admits Brexit lies; Protecting the web from its founders; Sanders x Hillary; Surveillance pricing v your dollars. Upcoming appearances: Philadelphia, Chicago, London, Edinburgh, Sydney, Melbourne, Brighton, London, South Bend. Recent appearanc
Pluralistic (Cory Doctorow) 35d ago Safety & alignmentPrivacy

How Google's Waymo is Scaling Robotaxis in 2026

The Race to Autonomous Driving is Heating Up🔥 A deepdive into Waymo. 🗺️🚘🛣️
AI Supremacy 35d ago Agents & autonomy

Policy (3)

Research (54)

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration

Reward models for Reinforcement Learning from Human Feedback (RLHF) pool preferences across thousands of annotators and fit one global affine calibrator, collapsing raters with systematically different rating-scale offsets and slopes into a single average-rater fit that does not match any individual annotator. PEBS is a per-rater empirical-Bayes shrinkage estimator: it fits per-rater affine calibrators on a held-out slice of each annotator's ratings and applies Morris-James-Stein empirical-Bayes
arXiv 35d ago

hia-gat: A Heterogeneous Interaction-Aware Graph Attention Network For Frame-Level Traffic Conflict Risk Prediction On Freeways

This paper formulates frame-level freeway risk assessment as a multi-agent scene graph-level binary classification problem, where each video or trajectory frame is labeled risky if any TTC- or PET-based conflict violates a specified severity threshold. We construct a relation-aware graph per frame with vehicles as nodes and two interaction types as edges: same-lane (longitudinal) and adjacent-lane (lateral), augmented with physics-informed edge features aligned to rear-end and lane-change confli
arXiv 35d ago Agents & autonomy

On the Inseparability of Instructions and Data in Shared-Embedding Sequence Models

Prompt injection is the top security risk for LLM-integrated applications, yet every defense proposed so far has been broken. We prove this is not a coincidence: in shared-embedding architectures that lack enforced control-data separation, perfect prompt-injection prevention is mathematically impossible. We formalize prompted systems as Prompted Action Models whose outputs include control-authoritative actions: refusal decisions, tool authorization, policy routing, and memory writes. We define S
arXiv 35d ago RegulationMilitary & security

Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction

Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations. Accurate prediction enables key downstream applications, such as advertising optimization and strategic content planning by users, creators, and platforms. Despite substantial progress, existing popularity prediction works often fail to jointly consider multimodal content and temporal social interaction signals. Moreover, the literature remains highly fragmented acro
arXiv 35d ago

Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching

Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are compared to determine whether they refer to the same real-world entity. Recent work has incorporated domain information and low-resource learning techniques to better adapt EM systems to realistic settings. While these approaches have demonstrated strong performance, it remains unclear how they behave under varying data constraints and levels of supervision in practice. In this pap
arXiv 35d ago Safety & alignment

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning

Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing complex tasks into executable actions. While small open source MLLMs are cost efficient and privacy preserving compared with commercial large models, they suffer from weak planning and limited cross website generalization. To address these limitations, we introduce the planning experience exploration and utilization (PEEU) method, which autonomously explores envir
arXiv 35d ago PrivacyAgents & autonomy

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

Text-to-image (T2I) diffusion models typically require substantial computational resources and cloud infrastructure, posing significant challenges for edge deployment in terms of latency, cost, and user privacy. We present JuZhou 1.0, an ultra-lightweight T2I foundation model designed for fully offline, on-device execution. JuZhou 1.0 achieves its efficiency through four key designs: (1) a compact image-generation backbone consisting of a 0.385B-parameter denoising U-Net and a 1.90M-parameter di
arXiv 35d ago Privacy

Prompt Injection in Automated Résumé Screening with Large Language Models: Single and Multi-Injection Settings

Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candidates to strategically manipulate algorithmic hiring systems. We study prompt injection in automated résumé screening, defined as subtle self-promotional text that introduces no new qualifications but is designed to influence LLM evaluations. Using controlled experiments, we show that prompt injection reliably improves applicant rankings when résumé quality is homogeneous and few ca
arXiv 35d ago Jobs & economy

From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Community Dynamics on 4chan

AI nudification uses generative models to create synthetic non-consensual sexually explicit imagery (SNEACI) of real individuals. Prior work has examined dedicated nudification platforms and model repositories, finding that most targets are female celebrities. However, the anonymous content community, where SNEACI is actively requested, generated, and exchanged, remains unexplored. In this work, we present a large-scale study of AI nudification in the wild, identifying 24,105 SNEACI items. We fi
arXiv 35d ago

A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO

We introduce the process harness, a new mechanism for uplifting legacy workflows into Agentic Business Process Management (Agentic BPM) without replacing the underlying workflow engine. A process harness places a policy-governed agentic layer around a deterministic workflow engine, intercepting designated control points to contribute reasoning, adaptation, and oversight while the engine retains structural authority over the process. To define the process harness rigorously, we develop the Task-D
arXiv 35d ago RegulationAgents & autonomy

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progress, and a few task-relevant future quantities, and those predictions drive advantage estimation, live
arXiv 35d ago Regulation

Joint Learning of Experiential Rules and Policies for Large Language Model Agents

For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model as natural-language rules for later prompting, or using trajectories and feedback to update the model parameters. The former is easy to interpret but can fall out of sync with the evolving policy; the latter improves the policy more broadly but provides only limited co
arXiv 35d ago RegulationAgents & autonomy

Efficient foundation decoders for fault-tolerant quantum computing

Foundation decoders, a class of high-capacity neural decoders, are leading candidates for fault-tolerant quantum computing, with accurate and efficient decoding at large code distances. However, their construction often faces a steep scaling barrier, as larger code distances rapidly amplify the cost of syndrome generation and neural optimization. To address this bottleneck, here we devise neural transfer unification (NTU), a unified framework for efficient foundation decoders. A central feature
arXiv 35d ago

Parametric Open Source Games

Open-source game theory studies agents whose behavior may depend on one another's decision procedures, but most existing models use discrete or symbolic programs. We introduce parametric open-source games, a continuous analogue of program equilibria in which players choose parameter vectors and semantics maps convert the full parameter profile into mixed actions in an underlying finite game. We establish equilibrium existence results, derive an exact coupling threshold at which selfish gradient
arXiv 35d ago Agents & autonomy

Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI)

Modern AI systems are increasingly deployed under non-stationary computational, demographic, and operational conditions in which static resource allocation strategies degrade both predictive performance and human-centric properties such as fairness and explainability. This paper presents AURORA-AI, an Adaptive Utility-driven Resource Orchestration framework for Resilient AI that unifies Hamilton-Jacobi-Bellman feedback control, Lyapunov-based stability monitoring, and a fairness-aware composite
arXiv 35d ago Bias & fairnessTransparency

Decision-Aligned Evaluation of Uncertainty Quantification

Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce decision-alignment, a criterion that reveals which evaluation metrics meaningfully align with downstream utilities. Applying this framework, we show that many widely used uncertainty metrics are either misaligned with common deci
arXiv 35d ago Safety & alignment

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirroring human psychological structure. We test the generality of these findings in two open-weight models, Apertus-8B-Instruct-2509 and Gemma-4-E4B-it, extracting emotion contrast vectors across all layers, using two model-generated corpora. We recover valence geometry for both models, with peak PC1--valence correlations
arXiv 35d ago

Compression-Driven Anomaly Detection in Brain MRI Using an Interpretable Quantum Autoencoder

We study a quantum autoencoder (QAE) for compression-driven anomaly detection in brain MRI data. The approach leverages angle encoding to map image patches into quantum states, followed by a variational encoder-decoder architecture trained to discard information via auxiliary trash qubits. Anomaly scores reflect the degree to which inputs resist compression relative to normal data, with higher scores corresponding to deviations from the learned normal manifold. Evaluated on publicly available br
arXiv 35d ago

Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions

Large language models (LLMs) are increasingly being integrated into mental health support tools and other psychologically sensitive conversational applications. In such settings, behavioral stability and consistency are important for trustworthy human-AI interaction. However, semantically similar concerns can be presented through different contextual framings, potentially eliciting different model responses. Such framing-sensitive variability may challenge user expectations regarding system beha
arXiv 35d ago HealthcareTransparency

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary instructions and compose high-level action plans, but they often generate motions that violate physical constraints. Physics-aware models improve realism through simulation or control, but they struggle with semantic com
arXiv 35d ago

XMSE-Aware Adaptive Empirical Bayes Estimation

Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter. This paper turns that diagnostic into a design principle. We propose an XMSE-aware mixed estimator that interpolates between ML and EB shrinkage. Its fixed-weight XMSE is a scalar qua
arXiv 35d ago Healthcare

A Deterministic Control Plane for LLM Coding Agents

LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitions, IDE-specific markdown -- is largely unmanaged. A prevalence study of 10,008 public GitHub repositories (n=6,145 agent config files) finds that agent configurations propagate as undeclared shared components: 10.1% of tracked paths are SHA-256 exact duplicates across independent repositories (fork-adjusted, threshold-independent), with 75.5% of clone pairs
arXiv 35d ago Agents & autonomy

Diagnosing Task Insensitivity in Language Agents

Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of this failure as task insensitivity: when faced with similar but distinct tasks, models might apply patterns learned during training and fail to solve the task at hand. We show that models often continue with actions aligned with the original task even when the instruction is semantically corrupted and cannot be directly answered. We further
arXiv 35d ago HealthcareAgents & autonomy

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollouts induces representation-space preference directions that sharply disagree with the batch majority, resulting in high-variance and destabilizing updates. We propose geoalign, a lightweight plug-in for rollout curation
arXiv 35d ago

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, we first conduct the systematic audit of ASR performance on real-world psychiatric interview data spanning Kannada, Hindi and Indian English, comparing eight state-of-the-art models including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini.
arXiv 35d ago HealthcareTransparency

Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization

Embedding-based retrieval ranks items by their similarity to a query in a shared vector space and usually aims to return the highest-scoring items. In many production settings this is not what is wanted: given a seed set that expresses a fine-grained pattern, one needs more items that both satisfy a target attribute and stay within that pattern. We formalize this as pattern-preserving attribute retrieval. The two goals pull against each other: averaging the seeds preserves the pattern but stays
arXiv 35d ago Regulation

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched. Existing vision-language CBMs often rely on pre-aligned encoders or global cosine similarity, which obscures fine-grained concept localization and fails to reflect true semantic geometry. In this work, we rethink concept alignment as a dynamic cross-modal transport pr
arXiv 35d ago Safety & alignmentTransparency

Fortress and Gatekeeper: Theorizing Transitive Trust in Third-Party Cybersecurity Risk Governance

Third-party vendors, such as analytics platforms, cloud services, identity providers, and software suppliers, are increasingly embedded in digital service delivery. While these arrangements enable scale and specialization, they also move customer data and security-relevant practices into environments that customers rarely see, select, or evaluate. This paper examines this problem through a document analysis of the November 2025 OpenAI-Mixpanel security incident. The incident serves as an illustr
arXiv 35d ago RegulationEnvironment

NaviCache: Test-Time Self-Calibration Caching for Video Generation

Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from calibration data dependency, prohibitive calibration duration, and susceptibility to distribution shifts, offline calibration-free methods eliminate these hurdles. However, since they rely on instantaneous zero-order approximations where the mapping between input and output differences varies in real-time, they are susceptible to observational noise and ignore th
arXiv 35d ago

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applications increasingly demand visually grounded commonsense inference and compositional reasoning, it remains unclear whether CLIP-style encoders can support such reasoning without architectural changes. To address this, we present ReasonCLIP-58M, a continual pretraining framework that integrates large-scale reasoning super
arXiv 35d ago Safety & alignment

AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing

Traditional dynamic pricing models in large-scale e-commerce suffer from limited interpretability, poor utilization of unstructured information, and misalignment with long-term business objectives such as cumulative Gross Merchandise Value (GMV), Return on Investment (ROI) and milestone achievement. We propose AIGP, a novel framework that leverages a Large Language Model (LLM) prompted with domain knowledge, structured data and textual context to make interpretable, knowledge-aware pricing decis
arXiv 35d ago Safety & alignmentFinance, VC & PE

ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration

The adoption of powerful diffusion models is hindered by their significant inference latency. Recent ``cache-then-forecast'' schemes alleviate this issue by accelerating DiTs using derivative-based polynomials, but they suffer from severe quality degradation at high acceleration ratios. Our analysis reveals its root cause: the discrete extrapolation performed on representations that are misaligned with the continuous diffusion trajectory and are numerically unstable. Thus, accelerated DiTs suffe
arXiv 35d ago Safety & alignment

Anatomy-Guided Residual Motion Diffusion for Controllable 4D Cardiac MRI Synthesis

Developing robust artificial intelligence models for 4D (3D + time) medical imaging is constrained by limited annotated data, inter-device domain shifts, and privacy restrictions. To address this, we propose a 4D controllable generative framework for anatomically consistent data augmentation. A semi-supervised variational autoencoder learns a compact latent representation of anatomical volumes while jointly predicting aligned segmentation masks in a unified framework. Anatomical structure is the
arXiv 35d ago PrivacyHealthcare

Robust Onion: Peeling Open Vocab Object Detectors Under Noise

The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis Robust Onion, an empirical study that uses controlled synthetic visual degradations to peel OV-ODs layer-by-layer, revealing how, why, and where robustness degrades, systematically analyzing feature collapse. Our findings reveal that models with similar vision backbones exhibit comparable robustness, driven by similar f
arXiv 35d ago

MLFFM-SegDiff: A Multi-Level Feature Fusion Diffusion Model for Skin Lesion Segmentation

Skin lesion segmentation is a key task in computer-aided dermatological diagnosis, where accuracy directly impacts downstream analysis and disease classification. However, dermoscopic images are challenging due to blurred boundaries, low contrast, large shape variations, and artifacts such as hair and shadows. Recently, diffusion models have shown strong performance in medical image segmentation thanks to their progressive denoising and distribution modeling capabilities. Nevertheless, existing
arXiv 35d ago Healthcare

Algorithmic Foundations of Deep Learning: Complexity-Theoretic Rates and a Characterization of Universal Approximation

Feedforward neural network (NN) expressivity is typically studied by emulating optimal basis-expansion schemes. While powerful, this perspective is incomplete: it primarily captures complexity through regularity, and therefore does not distinguish intuitively simple and complicated objects with comparable regularity, such as the square-root function and a typical Brownian path. The guiding message is that neural networks should be viewed not only as flexible basis functions, but also as models o
arXiv 35d ago

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization. This work presents NebulaExp, a fully transparent, ablation-driven post-training pipeline built on Qwen3-8B-base, covering two orthogonal model branches: general instruct model and complex reasoning-special
arXiv 35d ago Safety & alignmentTransparency

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization

Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos. While significant strides have been made in image stylization and video motion customization, simultaneously controlling multiple concepts, such as content, style, and motion, remains a major challenge. In this work, we systematically define the task of multi-concept video customization, which requires the joint control of content, style, and motion. To fac
arXiv 35d ago

Zero-Shot Size Transfer for Neural ODEs on Sparse Random Graphs: Graphon Limits and Adjoint Convergence

Graph Neural Differential Equations (GNDEs) model continuous-time graph dynamics by parameterizing Neural ODE velocity fields with Graph Neural Networks. Their local, size-independent filters suggest a zero-shot size-transfer principle: train on a small graph and deploy on larger, similar graphs without retraining. We develop a quantitative theory for this principle on sparse random graphs sampled from graphons. We consider Graphon Neural Differential Equations (Graphon-NDEs) and adjoint Graphon
arXiv 35d ago

LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction

Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. While existing predictors excel at minimizing standard displacement errors, they often overlook the adherence to lane topology of multimodal predictions, particularly for lower-probability modes. Consequently, predicted trajectories may violate physical and logical constraints, making the prediction set unreliable for safety-critical planning. In this paper, we
arXiv 35d ago Jobs & economy

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents

Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and act on a user's behalf. As they move from answering questions to operating over sensitive data, privacy becomes harder to enforce. An agent touches many data sources, runs multi-step workflows, keeps state across sessions, and acts with delegated permissions. Sensitive information can therefore leak not only through its final answer but through the queries it
arXiv 35d ago PrivacyAgents & autonomy

LLM-based Models for Detecting Emerging Topics in Service Feedback

Enhancing the analysis of service feedback is essential for public sector organizations, particularly tax administrations, where trust and compliance depend on fair and effective service delivery. As feedback volumes grow, identifying emerging service quality issues and potential disparities across diverse populations becomes increasingly challenging. Traditional approaches often rely on manual review or static expert-defined indicators, limiting scalability and the ability to capture complex pa
arXiv 35d ago Regulation

IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control

Complex multi-agent control tasks remain challenging for traditional rule-based and model-based approaches, motivating the adoption of learning-based methods. However, learning-based methods often struggle with sim-to-real transfer because they rely on accurate dynamics modeling or system identification and learn policies in low-level control spaces that are highly sensitive to dynamics mismatch, making them costly and fragile in complex environments. To address this issue, we propose a sim-to-r
arXiv 35d ago Safety & alignmentAgents & autonomy

Pingquanqi (Equalizer): A Cross-Domain Sociotechnical Framework for Human-Agent Interaction Governance

LLM agents are transitioning from experimental tools to permanent infrastructure -- a computational layer as enduring as the electrical grid. Like any infrastructure, they carry a cost chain from physical capital through enterprise investment to user consumption, ending at the user's most irreplaceable resource: lifetime. When unoptimized, this chain leaks, consuming user lifetime without adequate compensation. This paper proposes Pingquanqi (Equalizer), a cross-domain sociotechnical framework f
arXiv 35d ago RegulationAgents & autonomy

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need. Humanitarian organizations often lack the staff, time, and specialist expertise required to analyze this information at scale. Large language models (LLMs) may expand this capacity, but their reliability for coding qualitative humanitarian data has not been directly established. This benchmark study compares 46 LLMs to a huma
arXiv 35d ago

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP

Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning. Unlike traditional black-box QA, CRISP utilizes metric 3D Scene Graphs and an oracle intervention protocol to decouple latent reasoning capabilities from perceptual bottlenecks. This granular diagn
arXiv 35d ago Safety & alignmentHealthcare

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

The Abstraction and Reasoning Corpus (ARC) contains tasks that require summarizing patterns from limited grid samples and predicting output grids. Recently, many large language model based approaches have attempted to transform it into a text-based reasoning task. However, methods based on open-source models have generally yielded unsatisfactory results, while those relying on closed-source models are too costly. Current efforts mainly focus on data augmentation, constructing ARC-like data for m
arXiv 35d ago

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We show that conditioning a language or vision model on a narrow task suppresses its reporting of co-present, safety-critical signals it can otherwise report, a machine analogue of human inattentional blindness, produced by a different mechanism. Across radiology and driving text scenarios and chest-radiograph vision tasks, the ordinary focused instru
arXiv 35d ago Safety & alignment

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong, it spends more tokens than when it gets the same problem right; humans do the reverse, spending less time on the trials they get wrong. We separate two levels of deliberation: how response time tracks difficulty across items (registration), and, with item identity held fixed, whether an agent spends more on its own fail
arXiv 35d ago Agents & autonomy

Clinical Harness for Governable Medical AI Skill Ecosystems

Medical AI remains organized around isolated models, whereas care requires accountable capabilities that persist across time. We define clinical AI skills and propose the Clinical Harness, a runtime governance architecture that registers, orchestrates, constrains and monitors them. Using osteoporosis as an exemplar, we show how knowledge-driven, data-driven and physics-enhanced skills can support lifecycle care and provide a governed substrate for future medical agents.
arXiv 35d ago RegulationHealthcare

Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks

Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We study \textbf{retrieval-warmed energy-based reasoning (RW-EBR)} -- an IRED energy-based diffusion model \cite{du2024ired} augmented with a Modern Hopfield trajectory memory -- and contribute a \textbf{five-arm ablation methodology} (oracle, best-constant, per-query-random, shuffled, aligned) that separates three confounded effects: class-prior bias shift, stochas
arXiv 35d ago Bias & fairnessEnvironment

The Tilted Playing Field for Women in Science

Institutional prestige shapes access to resources, visibility, and collaboration opportunities in science. Yet whether prestige benefits researchers equally, and how it relates to differences in scientific productivity and collaboration, remains unclear. Here, we quantify prestige advantage as the relative likelihood that researchers at higher-ranked institutions have more collaborators and produce more high-impact papers compared to their lower-ranked peers. Analyzing nearly 5 million papers by
arXiv 35d ago Jobs & economy

AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns

AI healthcare chatbots are increasingly used to support health information seeking and self-management, yet their performance and impact on users remains to be studied. This study examines over 15,000 user reviews from 59 AI healthcare chatbot apps to explore how these systems function in everyday informational and emotional contexts. Topic modeling and interpretive analysis identify three recurring breakdowns: access barriers and service unreliability, user experience and interaction quality, a
arXiv cs.HC 35d ago Healthcare

Relationships in the age of AI: A review on the opportunities and risks of synthetic relationships to reduce loneliness

Loneliness is a pressing global health issue, yet traditional interventions often fall short due to scalability limitations and the individualized experiences of loneliness. The rise of generative artificial intelligence (AI) has enabled synthetic relationships (SRs)—ongoing associations with AI companions designed to simulate human-like social bonds. SRs offer, among other aspects, constant availability, adaptability, and emotional responsiveness, which potentially address loneliness. However,
OpenAlex 35d ago Healthcare