Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stop-feedback into dense per-step costs via return decomposition, then trains a constrained offline policy on the augmented dat
Record details
Published: 12 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Regulation
Retrieved: 13 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
arXiv cs.AI · 12 August 2026
When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide
arXiv cs.LG · 12 August 2026
Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT\&CK-Aligned Triage as a Worked Instance
arXiv cs.LG · 12 August 2026
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
arXiv cs.AI · 12 August 2026
Legally Mandated, but Still Inaccessible: Digital Tensions in Older Adults' Use of Norwegian Web Services
arXiv cs.HC · 12 August 2026
Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection
arXiv cs.LG · 12 August 2026
How to cite this record
ethics.ai (12 August 2026), “Redistribution-based Cost Inference Improves Sparse Safe Offline RL,” evidence record 19020, https://ethics.ai/record/19020 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.