Ethical and Technical Limits of Deepfake Speech Datasets
Claims about the robustness and fairness of deepfake speech detectors are only as credible as the datasets used to train and evaluate those systems. We present a dataset-level audit of the deepfake speech landscape. We compile and analyze 39 deepfake speech datasets, examining key attributes including accessibility, documentation, demographic and language coverage, dataset scale, and the underlying bona fide speech sources. Our audit reveals two important takeaways. Firstly, fairness assessment
Record details
Published: 9 June 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Misinformation · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Financial Audit Assistance using Misinformation Detection and Explanation
arXiv · 20 July 2026
What Do Deepfake Speech Detectors Actually Hear?
arXiv · 9 June 2026
Gender-based discrepancies in the algorithmic delivery of political ads on social media
arXiv · 9 June 2026
Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-\texorpdfstring{$δ$}{delta} Alignment
arXiv · 9 June 2026
Democracy in the Era of Artificial Intelligence
arXiv · 11 June 2026
Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard
arXiv · 7 June 2026
How to cite this record
ethics.ai (9 June 2026), “Ethical and Technical Limits of Deepfake Speech Datasets,” evidence record 1193, https://ethics.ai/record/1193 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.