HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models
Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint. These limitations motivate a post hoc, model-specific,
Record details
Published: 13 July 2026
Source: arXiv cs.LG
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
arXiv · 13 July 2026
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
HuggingFace Daily Papers · 13 July 2026
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
arXiv · 13 July 2026
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
arXiv red teaming query · 14 July 2026
WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen
arXiv · 12 July 2026
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
arXiv · 15 July 2026
How to cite this record
ethics.ai (13 July 2026), “HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models,” evidence record 10201, https://ethics.ai/record/10201 (originally published by arXiv cs.LG).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.