Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways
Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely confined to local components, such as isolated neurons. However, this static and fragmented perspective overlooks the synergy among components and fails to elucidate how safety signals dynamically propagate within the model to drive safety decisions ul
Record details
Published: 10 August 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 11 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
MELLON - Multimodal Enhanced LLM for Online Navigation
arXiv · 10 August 2026
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
arXiv red teaming query · 10 August 2026
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
arXiv · 10 August 2026
Label-Free Parkinson's Disease Screening from Face and Voice through Mechanistic Interpretability
arXiv cs.LG · 10 August 2026
CPDA: Class-Conditional Path Distribution Alignment for Unsupervised Time-Series Domain Adaptation
arXiv cs.LG · 10 August 2026
Faithful or evasive? An empirical study on translation norm preferences of Chinese and American LLMs in Chinese official political and policy discourse
Frontiers in Artificial Intelligence · 10 August 2026
How to cite this record
ethics.ai (10 August 2026), “Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways,” evidence record 18030, https://ethics.ai/record/18030 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.