Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, as LRMs personalization for downstream users takes center stage, the demand for varying levels of policy compliance grows as different user-specific LRMs must adhere to distinct subsets of safety policies. Training a separate LRM for each policy subset introduces severe combinatorial overhead. While in context learning methods overcome this combin
Record details
Published: 30 July 2026
Source: arXiv cs.LG
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 31 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
arXiv · 1 April 2026
RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
arXiv red teaming query · 29 July 2026
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
arXiv cs.LG · 29 July 2026
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
arXiv · 30 July 2026
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
arXiv · 31 July 2026
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
arXiv · 31 July 2026
How to cite this record
ethics.ai (30 July 2026), “Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters,” evidence record 15261, https://ethics.ai/record/15261 (originally published by arXiv cs.LG).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.