From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
Language models are commonly fine-tuned for safety alignment to refuse harmful prompts. One approach fine-tunes them to generate categorical refusal tokens that distinguish different refusal types before responding. In this work, we leverage a version of Llama 3 8B fine-tuned with these categorical refusal tokens to enable inference-time control over fine-grained refusal behavior, improving both safety and reliability. We show that refusal token fine-tuning induces separable, category-aligned di
Record details
Published: 9 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
arXiv · 11 March 2026
Mind the Gap: Pitfalls of LLM Alignment with Asian Public Opinion
arXiv · 6 March 2026
Large language models provide unsafe answers to patient-posed medical questions
OpenAlex · 13 February 2026
BayMOTH: Bayesian optiMizatiOn with meTa-lookahead -- a simple approacH
arXiv · 13 April 2026
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
arXiv · 27 April 2026
AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
arXiv · 18 May 2026
How to cite this record
ethics.ai (9 March 2026), “From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions,” evidence record 7520, https://ethics.ai/record/7520 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.