The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime
AI deployment in sensitive domains such as health care, credit, employment, and criminal justice is often treated as unsafe to authorize until model internals can be explained. This often leads to an excessive reliance on mechanistic interpretability to address a deployment challenge beyond its intended scope. We argue that the gate should instead be calibrated verification: authorization should be domain-scoped, independently checkable, monitored after release, accountable, contestable, and rev
Record details
Published: 11 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Jobs & economy · Healthcare · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
arXiv · 11 May 2026
Fairness of Explanations in Artificial Intelligence (AI): A Unifying Framework, Axioms, and Future Direction toward Responsible AI
arXiv · 11 May 2026
Do Vision Models Truly Forget? New Findings from Representation-Level Certification of Visual Unlearning in Vertical Federated Learning
arXiv · 19 May 2026
MILD: Mediator Agent System with Bidirectional Perception and Multi-Layered Alignment for Human-Vehicle Collaboration
arXiv · 2 May 2026
SAM-Sode: Towards Faithful Explanations for Tiny Bacteria Detection
arXiv · 20 May 2026
NEURON: A Neuro-symbolic System for Grounded Clinical Explainability
arXiv · 2 May 2026
How to cite this record
ethics.ai (11 May 2026), “The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime,” evidence record 4552, https://ethics.ai/record/4552 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.