Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning
Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as simple numerical inputs rather than aligning them with the performance of their policies. In this paper, we propose Q-ALIGN DT, a framework that enforces this alignment by ensuring the $Q$-value of the output policy is consistent with the input RTG. By leveraging a $Q$ function to provide dense guidance to CSMs and further fine-tuning it using an RT
Record details
Published: 27 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models
arXiv · 27 May 2026
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
arXiv · 27 May 2026
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
arXiv · 28 May 2026
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
arXiv · 27 May 2026
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
arXiv · 28 May 2026
AI Loss of Control Incident Management: Response & Resilience
arXiv · 28 May 2026
How to cite this record
ethics.ai (27 May 2026), “Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning,” evidence record 3562, https://ethics.ai/record/3562 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.