DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair
While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rejection of benign queries that merely appear risky. We present DDOR (Delta Debugging for OverRefusal), a fully automated and explainable framework for overrefusal testing and repair in a black-box setting, where only model inputs and outputs are accessible and internal safety mechanisms remain opaque. DDOR applies delta debugging to localize minimal
Record details
Published: 2 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard
arXiv · 7 June 2026
Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word Subsets
arXiv · 7 June 2026
Gram: Assessing sabotage propensities via automated alignment auditing
arXiv · 28 May 2026
Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback
arXiv · 27 May 2026
Agent Economics: An Entropy-Controlled Pluralistic Alignment Framework for Preventing Artificial Hivemind in Autonomous Agents
arXiv · 8 June 2026
From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constraints
arXiv · 27 May 2026
How to cite this record
ethics.ai (2 June 2026), “DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair,” evidence record 3213, https://ethics.ai/record/3213 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.