From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa
Low-resource African languages lack text corpora needed for language model training. We investigate whether ASR pipelines can extend text resources for two typologically distinct West African languages: Fongbe (tonal, diacritic-rich) and Hausa (non-tonal). We fine-tune MMS-300M on a curated 12.3-hour Fongbe dataset, achieving 9.48% WER on the ALFFA benchmark - a 78% relative reduction from the prior 44.04% baseline - while preserving tonal diacritics critical to the language. For Hausa, we apply
Record details
Published: 20 June 2026
Source: arXiv
Category: Research
Topics: Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa
arXiv · 24 June 2026
Neutralizing Structural Inequality in the Nigerian FinTech Sector
arXiv cs.CY · 14 July 2026
Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria
arXiv cs.AI · 6 August 2026
Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria
arXiv cs.CY · 7 August 2026
The Use of Learning Management Systems for Self-paced Learning: The Case at a South African Public Access Centre
arXiv cs.CY · 14 August 2026
Analysis Of Linguistic Stereotypes in Single and Multi-Agent Generative AI Architectures
arXiv · 19 March 2026
How to cite this record
ethics.ai (20 June 2026), “From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa,” evidence record 740, https://ethics.ai/record/740 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.