Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering
Open-weight video diffusion models can generate photorealistic unsafe content, from violence to misinformation, yet existing defenses either require expensive safety fine-tuning that degrades general capability, or apply external filters that are trivially bypassed by adversarial prompts. We present REINS (REpresentation-space INference-time Safety steering), a training-free method that aligns video diffusion models at inference time by steering their internal representations toward safe generat
Record details
Published: 15 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Misinformation
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Affective AI Safety: The Missing Piece in LLM Safety
arXiv · 22 June 2026
Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models
arXiv · 7 June 2026
Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard
arXiv · 7 June 2026
Amplifying the storm: Climate disinformation dynamics during natural disasters on right-wing extremist Telegram channels
HKS Misinformation Review · 23 July 2026
Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
HuggingFace Daily Papers · 6 August 2026
From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos
arXiv · 7 August 2026
How to cite this record
ethics.ai (15 June 2026), “Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering,” evidence record 935, https://ethics.ai/record/935 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.