Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward white-box attacks that directly identify and disrupt internal safety neurons or routes. However, existing safety defenses often rely on static safety units or fixed refusal pathways, leaving models highly vulnerable to targeted route-level white-box attacks. For that, we propose dynamic routing adaptive alignment (DRAA), a framework that introduces d
Record details
Published: 2 August 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment · Environment
Retrieved: 5 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
An AI-Based Decision-Support Pipeline for Day-Ahead Photovoltaic Forecasting
arXiv · 3 August 2026
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
arXiv red teaming query · 1 August 2026
AS-FedBridge: Pseudo-Spike Bridge Distillation for Heterogeneous ANN-SNN Federated Learning
arXiv cs.LG · 4 August 2026
How Usable Are Geospatial Foundation Models? A Systematic Evaluation of 89 Models
arXiv cs.HC · 4 August 2026
How Usable Are Geospatial Foundation Models? A Systematic Evaluation of 89 Models
arXiv cs.HC · 4 August 2026
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
HuggingFace Daily Papers · 31 July 2026
How to cite this record
ethics.ai (2 August 2026), “Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks,” evidence record 16570, https://ethics.ai/record/16570 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.