Leveraging RAG for Training-Free Alignment of LLMs
Large language model (LLM) alignment algorithms typically consist of post-training over preference pairs. While such algorithms are widely used to enable safety guardrails and align LLMs with general human preferences, we show that state-of-the-art alignment algorithms require significant computational resources while being far less capable of enabling refusal guardrails for recent agentic attacks. Thus, to improve refusal guardrails against such attacks without drastically increasing computatio
Record details
Published: 11 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Understanding the Effects of Safety Unalignment on Large Language Models
arXiv · 2 April 2026
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
arXiv · 12 May 2026
Guided Streaming Stochastic Interpolant Policy
arXiv · 11 May 2026
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
arXiv · 12 May 2026
No More, No Less: Task Alignment in Terminal Agents
arXiv · 12 May 2026
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
arXiv · 13 May 2026
How to cite this record
ethics.ai (11 May 2026), “Leveraging RAG for Training-Free Alignment of LLMs,” evidence record 4531, https://ethics.ai/record/4531 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.