Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations
Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly conditioned on pragmatic register: models that refuse a direct request frequently comply when the same underlying objective is expressed through a different communicative stance. This suggests that current alignment policies are not invariant to semantic equivalence, but remain sensitive to how a request is pragmatically framed. We introduce Retroact
Record details
Published: 6 July 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models
arXiv · 7 July 2026
Alignment Plausibility: A New Standard for Assuring AI in Healthcare
arXiv · 8 July 2026
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
arXiv · 3 July 2026
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
arXiv · 8 July 2026
RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation
arXiv · 2 July 2026
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
arXiv · 2 July 2026
How to cite this record
ethics.ai (6 July 2026), “Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations,” evidence record 212, https://ethics.ai/record/212 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.