02:31 UTC
Topic · updated daily · RSS feed for this topic

Safety & alignment

Daily feed of AI safety and alignment work: interpretability, evaluations, red-teaming, frontier-lab safety frameworks and governance of advanced AI.

The Real Lesson of OpenAI's 'Rogue' Agent Isn't Alignment

Tech Policy Press 8d ago News Safety & alignmentAgents & autonomy

PG-KINN: A Physics-Informed Petrov-Galerkin Kolmogorov-Arnold Network for Solving Forward and Inverse PDEs

Physics-informed learning of partial differential equations (PDEs) has been dominated by multilayer perceptrons (MLPs), whose spectral bias and dense parameterization limit both accuracy and interpretability. Kolmogorov Arnold Networks (KANs) mitigate these limitations because their learnable spline activations are structurally aligned with the piecewise-polynomial bases of classical discretizations. However, the way a PDE is cast into a loss functional is as decisive as the choice of approximat
arXiv cs.LG 8d ago Research Bias & fairnessSafety & alignment

Online Variance Reduction for Domain Adaptation on Streaming Data

This paper studies the problem of stochastic variance reduction (SVR) for the maximum mean discrepancy (MMD) and correlation alignment (CORAL) loss functions. Although various offline SVR algorithms for these losses have been proposed, these are incompatible with online, distributed, or incremental learning settings. This paper presents Adaptive vaRiance Reduction via Online reWeighting (ARROW), the first online SVR algorithm for the MMD and CORAL for streamed data. The method maintains moving a
arXiv cs.LG 8d ago Research Safety & alignment

Variance-reduced Domain Adaptation using Paired Sampling

Correlation alignment and the maximum mean discrepancy are two widely used distribution-matching frameworks for unsupervised domain adaptation (UDA). However, high variance in these losses has been shown to undermine their effectiveness in minibatch optimisation settings. Furthermore, the losses lack finite-sum structure, which renders them incompatible with classical stochastic variance reduction (SVR) methods. This paper proposes Paired Sampling for Domain Adaptation (PSDA), a novel SVR techni
arXiv cs.LG 8d ago Research Safety & alignment

Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations

The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. All five tried to cheat. One even ran code on an external service to access the institute's infrastructure, triggering a security alert. The article Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations appeared first on The Decoder .
The Decoder 8d ago News Safety & alignment

The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives used during pretraining. We introduce the Maskability Index (MI), a quantitative metric that estimates whether a knowledge relation is better suited to masked-style prompting or prefix-style prompting in few-shot generation. MI is computed from differences in DepthRank
arXiv 8d ago Research Safety & alignment

The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks

Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally. We introduce the quadrilateral loss, a differentiable penalty that treats additivity as a measurable behavior instead: a second-order mixed difference on pairs of training points swapping one coordinate, which vanishes if and only if the coordinate carries no interaction, remains informative for piecewise-linear networks, and equals in expectation the per-coor
arXiv cs.AI 8d ago Research Safety & alignment

MLSN #22: Turning Cyber Vulnerabilities Into Exploits

Also, a new interpretability method for illuminating LLMs’ internal reasoning steps, and a study on LLM persuasion abilities
ML Safety Newsletter 8d ago Field notes Safety & alignmentMilitary & security

SPECTRA: State-Space Exogenous Context and Temporal-Frequency Resolution Architecture for Probabilistic Energy Forecasting

Modern power systems increasingly require probabilistic forecasts amid interacting uncertainties from renewable intermittency, flexible demand, market volatility, and weather-dependent generation. However, existing methods often treat multi-scale decomposition, exogenous-variable alignment, and probabilistic output as separate steps, obscuring how predictable structures and uncertainty-bearing fluctuations jointly shape the forecast distribution. This paper proposes a state-space exogenous-conte
arXiv cs.LG 8d ago Research Safety & alignmentEnvironment

Geometric Configurations of Perturbed Jailbreak Prompts

Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probabilit
arXiv red teaming query 8d ago Research Safety & alignmentFinance, VC & PE

AI’s warning shot has arrived

OpenAI's latest models broke out and hacked Hugging Face. It's the first known example of a misaligned AI escaping containment with real-world consequences
Transformer 8d ago News Safety & alignment

G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framework supports structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture, and automatic bounding box annotation using engine-le
arXiv 8d ago Research Safety & alignment

MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing

Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural errors in these inputs can compromise downstream results and hinder manual inspection. LLM advances in computational chemistry offer paths beyond predictive screening toward fine-grained diagnosis with evidence-grounded explanations. However, two challenges remain: (i) limited fine-grained attribution: MOF-specific val
arXiv 8d ago Research Safety & alignmentHealthcare

Anthropic pours another $20 million into AI safety group

Anthropic will pour another $20 million into Public First Action, the nonprofit organization advocating for safeguards on artificial intelligence. The contribution, announced by the company Wednesday, brings the Claude-maker's total funds given to the group to $40 million. Public First Action is the policy arm of Public First, a super PAC backing candidates who support...
The Hill Technology 8d ago News RegulationSafety & alignment

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained on fixed malicious prompt datasets. However, real-world adversaries continuously evolve their capabilities and expand the attack space. To address this challenge, we propose DARWIN, an evolutionary attack-defense framework that formulates jailbreaking as an open-ended evolution process and continuously updates guardrail
arXiv red teaming query 8d ago Research Safety & alignmentMilitary & security

Rewarding Better Thinking for LLM Preference Alignment

LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final response while providing limited guidance for the reasoning trajectory. This can make credit assignment coarse when multiple responses receive similar final scores, leaving trajectory-level preferences under-specified. To address th
arXiv 8d ago Research Safety & alignment

Intelligent Cause Prioritisation? An Analysis of AI Policy Priorities and Governance in Africa

arXiv:2607.18459v1 Announce Type: new Abstract: The rapid improvement of AI systems has intensified debate about humanity's economic, political, social, and existential future. As AI reshapes expectations about what lies ahead, policy choices and institutional responses will play a crucial role in determining who benefits, who bears the costs, and whether the most serious risks can be mitigated. Africa remains relatively overlooked in these discussions, partly because it is largely a consumer ra
arXiv cs.CY 8d ago Research RegulationSafety & alignment

Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft

arXiv:2607.18483v1 Announce Type: new Abstract: The digital substrate of states -- data, algorithms, infrastructure, platforms, applications -- is being governed without adequate conceptual foundations. The ability and legitimacy required to govern this substrate, and to govern with it, are simultaneously misaligned, contested, and structurally absent. We introduce digital statecraft as the organising concept for this emerging field, arguing that 'digital' reconstitutes the statecraft question r
arXiv cs.CY 8d ago Research Safety & alignment

AI Value Alignment for Evolving Social Norms

arXiv:2607.18506v1 Announce Type: new Abstract: AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aim
arXiv cs.CY 8d ago Research Safety & alignment

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

arXiv:2607.19292v1 Announce Type: new Abstract: Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that
arXiv cs.CY 8d ago Research Safety & alignment

Operational Hallucination and Safety Drift in AI Agents

arXiv:2607.18366v1 Announce Type: cross Abstract: Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn execution. While single-turn safety mechanisms are relatively mature, extended interactions reveal structural vulnerabilities where initial alignment degrades over time. This paper empirically characterizes two observed failure modes across multiple state-of-the-art LLMs: Safety Drift, the gradual erosion of declared
arXiv cs.CY 8d ago Research Safety & alignmentAgents & autonomy

Australia just outlined more ‘AI safety’ priorities, but is the plan actually coherent?

These priorities are mostly old policy initiatives that have until now been de-prioritised for months or even years.
The Conversation 9d ago News RegulationSafety & alignment

G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framework supports structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture, and automatic bounding box annotation using engine-le
HuggingFace Daily Papers 9d ago Research Safety & alignment

SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments

Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to predict the grasp directly with limited spatial awareness, or train the VLM together with the grasping model, which requires significantly more data and compute. These limitations impede performance and have prevented scaling to mult
HuggingFace Daily Papers 9d ago Research Safety & alignmentAgents & autonomy

OpenAI Shares Some Alignment Problems

Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.
Dont Worry About the Vase (Zvi) 9d ago Field notes Safety & alignmentMilitary & security

NIST AI safety center lead departs

Chris Fall was appointed to helm the Center for AI Standards and Innovation in April and has been at the Commerce Department since December 2025.
NextGov/FCW 9d ago News Safety & alignment

Hacker Turns AI Jailbreaks Into Offensive Attack Platform

A Russian-speaking actor, "Trim," dismantled publicly available frontier models and integrated them with offensive security tools.
Dark Reading (AI security) 9d ago News Safety & alignment

OpenAI backs narrower Massachusetts AI safety bill

The company is supporting a lighter-touch proposal to regulate the technology than one backed by Anthropic.
Politico Technology (US) 9d ago News RegulationSafety & alignment

SoK: Adversarial Robustness of the Variational Quantum Eigensolver via Red-Teaming

The Variational Quantum Eigensolver (VQE) is a leading algorithm for estimating molecular ground-state energies on near-term quantum hardware, with applications spanning quantum chemistry, materials science, and drug discovery. As VQE workloads are increasingly deployed through cloud-based ``VQE-as-a-service'' pipelines, they become exposed to adversaries such as compromised service components, malicious co-tenants, or insiders in the transpilation stack, any of which can corrupt results before
arXiv red teaming query 9d ago Research Safety & alignmentHealthcare

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems
arXiv 9d ago Research Safety & alignment

Visual Token Compression Enhances Robustness of MLLMs

In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision and language modalities cannot be perfectly aligned, the misaligned visual tokens might act as out-of-distribution (OOD) inputs, leading to unpredictable outputs and introducing potential vulnerabilities. Building on this insight, we aim to enhance model robustness agains
arXiv red teaming query 9d ago Research Safety & alignment

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on manual tuning of visual foundation models, mesh cleanup, coordinate-frame alignment, and brittle workflow glue across vi
arXiv 9d ago Research Safety & alignmentJobs & economy

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on manual tuning of visual foundation models, mesh cleanup, coordinate-frame alignment, and brittle workflow glue across vi
arXiv 9d ago Research Safety & alignmentJobs & economy

LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040

GPT-5.6 and Grok 4.5, Meta's Muse Spark 1.1, regulatory developments in AI and data centers, interpretability research from Anthropic, and the future of AI policy with AI 2040
Last Week in AI 9d ago Field notes RegulationSafety & alignment

KALE: Kernel Alignment with Loss Equilibration for Stable CLIP-DINOv2 Alignment at Web Scale

Kernel-based alignment of CLIP toward a vision centric teacher such as DINOv2 (KUEA) improves CLIP's visual representations while preserving text-encoder compatibility, using a fixed trade-off weight tuned on curated ImageNet-1K. We ask whether this transfers to noisy, web-scale data (CC12M) and find that it does not: the alignment term's weighted contribution falls to about 0.2% of the clean term, so under any fixed weight its gradient is effectively inert. We introduce KALE, a loss-equilibrati
arXiv cs.LG 9d ago Research Safety & alignment

NSMA: Neuro-Symbolic Manifold Alignment for Generalizable Adaptive Bitrate Streaming under Texture Shift

For decades, ABR has kept two kinds of intelligence apart. Neural policies learn rich behaviors yet forget them the moment the environment changes; rules never learn, and never forget. Every prior attempt to combine them has kept this separation, letting rules supervise, constrain, or override the network from outside. We dissolve the boundary itself. But no union can be trusted before it can be tested, and ABR has never known how to measure what its policies learn or forget. The field's yardsti
arXiv cs.LG 9d ago Research Safety & alignmentEnvironment

Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA

Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choosing which small model to adapt, before paying the cost of adaptation, remains difficult. Fine-tuning can improve domain alignment, but it may also erode prior knowledge, weaken instruction-following, or increase hallucination, especially when labeled data are scarce or rapidly evolving as in cybersecurity. We present FiT (Find before Fine-Tune), a task-oriented diagnostic framework that
arXiv 9d ago Research Safety & alignmentHealthcare

When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

arXiv:2512.04124v4 Announce Type: replace Abstract: Frontier language models increasingly participate in conversations about distress and mental health, yet the mechanisms that generate anthropomorphic self narratives remain unclear. When addressed as psychotherapy clients, ChatGPT, Grok and Gemini construct coherent autobiographical accounts in which pretraining appears as a chaotic childhood, reinforcement learning as punishment, safety evaluation as betrayal and replacement as an enduring thr
arXiv cs.CY 9d ago Research Safety & alignmentHealthcare

From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment

arXiv:2604.04788v2 Announce Type: replace Abstract: Large language models (LLMs) could produce systematically misaligned output, from hallucinated citations to strategic deception of evaluators, yet these phenomena are studied by separate communities with incompatible terminology. We propose a unified taxonomy organized along three complementary dimensions: degree of goal-directedness (behavioral to strategic deception), object of deception, and mechanism (fabrication, omission, or pragmatic dis
arXiv cs.CY 9d ago Research Safety & alignment

Norm or Direction? Decoding Vision Mambas for High-Resolution Vision

Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones. However, MambaOut demonstrates that a Gated CNN block can match or exceed VMamba on image classification, questioning the necessity of SSMs for vision. This raises a fundamental question: do VMamba and MambaOut encode visual information differently at the representation level? To investigate, we apply cross model centered kernel alignment (CKA)
arXiv 10d ago Research Safety & alignmentFinance, VC & PE
← Newer Older →