HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models
Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. Such static designs struggle to maintain a cross-category safety boundary while generating constructive responses tailored to specific risks and avoiding over-refusal of benign inputs. To address these limitations, we propose HiRoute, an input-adaptive hierarchi
Record details
Published: 13 August 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment
Retrieved: 14 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Philosophical vertigo with artificial intelligence
arXiv cs.CY · 13 August 2026
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
arXiv · 13 August 2026
Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging
arXiv · 13 August 2026
From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion
arXiv · 13 August 2026
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds
arXiv · 13 August 2026
ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning
arXiv cs.LG · 13 August 2026
How to cite this record
ethics.ai (13 August 2026), “HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models,” evidence record 19472, https://ethics.ai/record/19472 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.