01:25 UTC
Topic · updated daily · RSS feed for this topic

Safety & alignment

Daily feed of AI safety and alignment work: interpretability, evaluations, red-teaming, frontier-lab safety frameworks and governance of advanced AI.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cr
HuggingFace Daily Papers 5d ago Research Safety & alignment

AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines

Outside safety experts say the models behind this week's hack may have crossed OpenAI's own 'critical' risk line, something that would require the company to halt development.
Fortune AI 5d ago News Safety & alignment

Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation

Gene regulatory network modeling often requires balancing predictive accuracy and mechanistic interpretability. In this work, we compare continuous surrogate models and a discrete mechanistic model on the same \textit{Arabidopsis thaliana} induced systemic resistance (ISR) dataset, using both the raw continuous gene-expression measurements and their sign-binarized representation. The study considers eight defense-related genes measured over nine time points and evaluates two continuous predictor
arXiv cs.LG 5d ago Research RegulationSafety & alignment

Quoting Boris Cherny

More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. — Boris Cherny , here's that System Card section , page 73 Tags: prompt-injection , anthropic , claude , generative-ai , ai , llms , boris-cherny
Simon Willisons Weblog 6d ago Field notes Safety & alignment

PRIMS: Physics-guided Representation for Fluid Identification in Multimodal Sensing

Accurate on-device fluid identification is essential for microfluidic applications, yet maintaining reliability under varying flow, pressure, and temperature remains a key challenge. Existing learning-based methods often treat sensor signals as domain-agnostic features, neglecting the underlying physical relationships that govern fluid behavior, thereby limiting generalization and interpretability. To address this, we propose PRIMS, a physics-aware multimodal Transformer that integrates physical
arXiv 6d ago Research Safety & alignment

Interior interpretability with attention rollout: contraction and propagation profiles in Transformers

Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretability}, a propagation-based perspective on internal model organization, and instantiate it for tabular Transformers using attention rollout. We interpret rollout as a row-stochastic operator encoding attention-mediated propagation between feature tok
arXiv cs.AI 6d ago Research Safety & alignment

A Roadmap to Impactful Pluralistic Alignment Research

Pluralistic value alignment---the goal of building AI systems that represent and serve diverse human values and perspectives---has emerged as an active research agenda. Yet, there's no public evidence that it has shaped the training or evaluation of the AI systems people actually use. We audit the public behavior documents and evaluations of frontier labs, finding none name pluralism as a goal, and as of this writing, no clear indication that production models are explicitly trained or tested fo
arXiv 6d ago Research Safety & alignmentTransparency

Huawei’s Ren Zhengfei backs Tau Scaling Law to beat US sanctions

Huawei Technologies founder Ren Zhengfei has thrown his weight behind the Chinese tech giant’s Tau Scaling Law, framing the new chip-design principle as an existential necessity to survive tightening US sanctions. “The Tau Scaling Law is the only path forward for Huawei to break the siege,” Ren said in recent internal remarks posted on Huawei’s Xinsheng Community intranet, according to a report on Friday by the state-backed Science and Technology Daily. Unlike companies with other technological.
SCMP Tech (HK/CN) 6d ago News RegulationSafety & alignment

Autoregressive EHR Foundation Models with Multimodal Inputs

Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way. We present a framework for conditioning such models on auxiliary clinical modalities, including ECG waveforms, chest X-ray images, and clinical notes, using modality-specific latent compression and gated cross-attention with temporal alignment. We investig
arXiv cs.LG 6d ago Research Safety & alignmentHealthcare

AI success depends on alignment

Research shows AI delivers real business outcomes only when organizations align their data, governance, and priorities across teams—not just when they adopt new tools.
Fast Company 6d ago News RegulationSafety & alignment

AI success depends on alignment

Research shows AI delivers real business outcomes only when organizations align their data, governance, and priorities across teams—not just when they adopt new tools.
Fast Company 6d ago News RegulationSafety & alignment

AI success depends on alignment

Research shows AI delivers real business outcomes only when organizations align their data, governance, and priorities across teams—not just when they adopt new tools.
Fast Company 6d ago News RegulationSafety & alignment

AI success depends on alignment

Research shows AI delivers real business outcomes only when organizations align their data, governance, and priorities across teams—not just when they adopt new tools.
Fast Company 6d ago News RegulationSafety & alignment

AI success depends on alignment

Research shows AI delivers real business outcomes only when organizations align their data, governance, and priorities across teams—not just when they adopt new tools.
Fast Company 6d ago News RegulationSafety & alignment

Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human
arXiv 6d ago Research Safety & alignment

dRAE: Representation Autoencoder with Hyper-Spherical Codes

In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales
arXiv cs.AI 6d ago Research Safety & alignment

Europe's Multilingual Reality Exposes AI Security Gaps

The AI security layer and guardrails for many AI products don't evenly protect against jailbreaking and unsafe actions in every single language.
Dark Reading (AI security) 6d ago News Safety & alignment

dRAE: Representation Autoencoder with Hyper-Spherical Codes

In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales
HuggingFace Daily Papers 6d ago Research Safety & alignment

TextSLIP: Text Self-Supervised CLIP for Medical Report Generation

Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pretraining (CLIP) has advanced medical vision language modeling, existing CLIP-style approaches may still provide insufficient fine-grained semantic supervision for complex report generation. Standard CLIP primarily optimizes cross-modal alignment, without explicitly structuring the textual embedding space that guides visual representation learning.
arXiv 6d ago Research Safety & alignmentHealthcare

Hong Kong must wake up to the cold hard geopolitics of AI

Officials talking about artificial intelligence (AI) often speak in terms of the next industrial revolution or existential risks. Such rhetoric ultimately does little to capture how AI filters into the ordinary workings of a city like Hong Kong. The geopolitics of AI are no longer abstract. They are beginning to shape who can use which tools – and on what terms. Over the past few months, two major international banks in the city have reportedly felt the impact. Anthropic’s Claude, a generative..
SCMP Tech (HK/CN) 6d ago News Safety & alignment

Adversarial Prompts for Acceptance Collapse in Speculative Decoding

Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence masks a severe operational vulnerability: draft-target alignment can be systematically attacked. In this paper, we introduce ADSD, which, to the best of our knowledge, is the first prompt-suffix attack that collapses verifier acceptance by pushing draft probability mass t
arXiv cs.LG 7d ago Research Safety & alignment

LAMAR: An Open Language-Aware Multilingual Alignment Reranker

In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple languages, which are subsequently reranked before answer generation. However, it remains unclear whether existing multilingual rerankers consider document language when ordering semantically relevant candidates. Our analysis shows that these rerankers do not consistently prioritize documents written in the same language as the query when semantically equivalent documents are available
HuggingFace Daily Papers 7d ago Research Safety & alignment

Spectral Prior for Reducing Exposure Bias in Diffusion Models

Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequency-dependent discrepancies between training and inference, which can be interpreted as frequency-dependent SNR error. Crucially, the direction of this mismatch varies across models and timesteps, indicating that fixed correction rules do not generalize. We propose Spectral Alignment (SPA), a lightweight, guidance-based method that calibrates the
HuggingFace Daily Papers 7d ago Research Bias & fairnessSafety & alignment

Projection Pursuit CPCANet for Domain Generalization

Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment methods, such as CPCANet, extract domain-invariant structures through batch-wise Common Principal Component Analysis (CPCA). However, CPCANet suffers from rank-deficient covariance estimation due to the small-sample-size issue in mini-batch training. To address this limitation, we propose Projection Pursuit CPCANet (PP-CPCANet), a covariance-free framework that learns a global ortho
HuggingFace Daily Papers 7d ago Research Safety & alignment

What AI Red-Team Evaluations Can and Cannot Prove

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated e
arXiv red teaming query 7d ago Research Safety & alignment

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training, the student generates reasoning trajectories conditione
arXiv cs.LG 7d ago Research RegulationSafety & alignment

DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation

Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised model DINOv3 provides strong structured visual representations, its lack of native textual alignment hinders its direct application to OVSS. To bridge this gap, we propose DINOde, an ODE-based framework that continuously aligns CLIP text embeddings with the DINO visual manifold. Our approach employs two complementary components: (i) Semantic Text Flo
arXiv 7d ago Research Safety & alignment

Emergent Misalignment Recruits a Pre-existing Persona Subspace

Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment. We ask why the narrow lesson generalizes at all, and we find that narrow fine-tuning recruits a persona structure that is present in the model before the fine-tune exists. From a frozen instruction-tuned model (Qwen2.5-14B-Instruct) we extract per-domain persona subspaces by contrastive teacher forcing and fi
arXiv cs.LG 7d ago Research Safety & alignment

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA's pseudo label for text-independent alignment (Chengdu-F
arXiv 7d ago Research Safety & alignment

AI #178: A Fire Alarm For General Intelligence

The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.
Dont Worry About the Vase (Zvi) 7d ago Field notes Safety & alignmentAgents & autonomy

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired with explicitly harmful queries. To understand the underlying mechanism of this vulnerability, we present V-DEAL, a three-level diagnostic framework that jointly analyzes this failure across model behaviour, understanding,
arXiv 7d ago Research Safety & alignmentHealthcare

The robot byline is quietly disappearing

The most obvious use of generative AI is writing. It’s right there in the name—large language models (LLMs) are all about reading, organizing, analyzing, and conjuring words—which is exactly why many in the journalism profession have been going through a kind of existential crisis these past few years. And the crisis isn’t just theoretical. As artificial intelligence systems get better at writing, a growing number of newsrooms are using AI to help not just with analysis, process, and ideas, but
Fast Company Tech 7d ago News Safety & alignmentAgents & autonomy

Amplifying the storm: Climate disinformation dynamics during natural disasters on right-wing extremist Telegram channels

Climate change amplifies natural disasters, posing an existential threat to our society. However, a digital storm is raging on right-wing extremist Telegram channels where climate disinformation works to delegitimize scientific consensus. Investigating the factors that amplify climate disinformation is as critical to combating it as understanding natural disasters and their drivers. We The post Amplifying the storm: Climate disinformation dynamics during natural disasters on right-wing extremist
HKS Misinformation Review 7d ago Research Safety & alignmentMisinformation

Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?

Yes, but less than had they been schemers.
Redwood Research 7d ago Field notes Safety & alignment

Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?

Alignment Forum 7d ago Research Safety & alignment

TwistedMerge: Certified Higher-Order Diagnostics and Abstention for Model Merging

Model merging combines independently trained or fine-tuned models, but pairwise alignability does not imply globally consistent alignment. We formulate merging as a finite descent problem in which checkpoints are local objects, alignment maps are transitions, and cycle products are residuals. TwistedMerge is a conservative certification pipeline that separates fixed-chart averaging, synchronization-removable gauge inconsistency, a certified central obstruction on a specified comparison complex,
arXiv 7d ago Research Safety & alignmentHealthcare

Code Monitor Red Teaming for Public-Test-Passing Code

Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has passed public tests, can a weaker LLM verifier identify the residual hidden bugs? We introduce Code Monitor Red Teaming, a monitor-red-teaming protocol that fixes a public-check information boundary while varying generator pressure, verifier scaffolding, and weak-to-strong capability. We instantiate it as CodeMonitorBen
arXiv red teaming query 7d ago Research Safety & alignment

Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning models remains con- strained by limited interpretability and the hallucination risk of large language models (LLMs). Existing CNN+Grad-CAM+multimodal LLM frameworks can generate ECG reports, but their explanations are often only weakly grounded in established diagnostic criteria, reducing trust- worthiness and reproducibility. We propose a guide-grounded multimodal framework that explic
arXiv 8d ago Research Safety & alignmentHealthcare

OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened

This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break in to Hugging Face, all so it could cheat on the test by stealing the answers. Along the way it helped make the strongest case yet for how the imbalance of model availability is hurting our ability to secure our software. Here's what hap
Simon Willisons Weblog 8d ago Field notes Safety & alignment

Language Models Embody and Amplify Human Cognitive Distortions: What Is to Be Done?

Human judgment is fundamentally prone to error. A promise of AI is that it will rid decisions of bias and ensure a fairer and safer world for all. Yet research unequivocally demonstrates that LLMs exhibit consequential sociocognitive biases. We alert readers that bias in AI (a) is covert and ironically a feature of alignment goals, (b) is not merely a mirror, but an amplifier of human bias, (c) intensifies across model generations, and (d) even transmits bias to humans. Given the potentially sei
arXiv 8d ago Research Bias & fairnessSafety & alignment
← Newer Older →