Principles and Guidelines for Randomized Controlled Trials in AI Evaluation
arXiv:2605.02050v2 Announce Type: replace Abstract: This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established practices from disciplines with established RCT traditions, including software engineering, economics, clinical and health sciences, and psychology, we synthesize five principles drawn from established validity frameworks and open-science standards on transparency, repeatability, and verification, which together
Record details
Published: 28 July 2026
Source: arXiv cs.CY
Category: Research
Topics: Healthcare · Transparency
Retrieved: 28 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Fairness Interventions in Classification: A Study on AI Explainability
arXiv cs.CY · 28 July 2026
Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks
arXiv cs.CY · 28 July 2026
KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability
arXiv cs.AI · 27 July 2026
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
arXiv cs.AI · 28 July 2026
Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis
arXiv cs.AI · 27 July 2026
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
arXiv cs.AI · 27 July 2026
How to cite this record
ethics.ai (28 July 2026), “Principles and Guidelines for Randomized Controlled Trials in AI Evaluation,” evidence record 13811, https://ethics.ai/record/13811 (originally published by arXiv cs.CY).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.