Each supported language is joined to the africa-corpus by canonical ISO 639-3, and real sentences are run through the engine. Coverage = share of alphabetic characters mapped to a known unit (the rest fall through as unknown). Generated by scripts/corpus_qa.py.
| Bucket | Languages |
|---|---|
| Excellent (≥95%) | 177 |
| Good (85–95%) | 18 |
| Partial (60–85%) | 8 |
| Poor (<60%) | 4 |
| Tested on corpus | 207 |
Languages without a corpus match are supported but not verified this way.
| Code | Coverage | Language |
|---|---|---|
ttq |
33% | Tawallammat Tamajaq |
hau-niger |
33% | Hausa |
sus |
50% | Susu |
shu |
60% | Arabic (Chadian) |
ney |
73% | Neyo |
bjo |
74% | Banda |
nwb |
77% | Nyabwa |
mfq |
80% | Moba |
shk |
82% | Shilluk |
azo |
82% | Awing |
mfb |
84% | Mbembe |
pny |
85% | Pinyin |