Skip to content

Latest commit

 

History

History
30 lines (25 loc) · 1.04 KB

File metadata and controls

30 lines (25 loc) · 1.04 KB

Phonemisation coverage (real corpus text)

Each supported language is joined to the africa-corpus by canonical ISO 639-3, and real sentences are run through the engine. Coverage = share of alphabetic characters mapped to a known unit (the rest fall through as unknown). Generated by scripts/corpus_qa.py.

Bucket Languages
Excellent (≥95%) 177
Good (85–95%) 18
Partial (60–85%) 8
Poor (<60%) 4
Tested on corpus 207

Languages without a corpus match are supported but not verified this way.

Below 85% — review (mostly script mismatches: corpus text in a different script)

Code Coverage Language
ttq 33% Tawallammat Tamajaq
hau-niger 33% Hausa
sus 50% Susu
shu 60% Arabic (Chadian)
ney 73% Neyo
bjo 74% Banda
nwb 77% Nyabwa
mfq 80% Moba
shk 82% Shilluk
azo 82% Awing
mfb 84% Mbembe
pny 85% Pinyin