There is a map you can draw of the world’s languages sorted by how well they are served by modern language technology. English sits at one end with tens of thousands of hours of training data behind it. At the other end are several thousand languages with effectively nothing.

The Horn of Africa sits close to the empty end, and it is worth being specific about what that means.

The languages

Somali — roughly 25 million speakers across Somalia, Somaliland, Djibouti, Ethiopia, and Kenya, plus a substantial global diaspora. Minimal training data, poor performance from every system we have tested.

Amharic — over 30 million first-language speakers and a working language of the Ethiopian federal government. Better served than most languages in the region, with some academic datasets and partial commercial coverage, but still far below what its speaker numbers would suggest.

Oromo — comparable to Amharic in speaker numbers, considerably worse in coverage. Very limited data, minimal commercial support.

Tigrinya — roughly seven million speakers across Eritrea and northern Ethiopia. Almost no speech technology coverage.

Afar — between two and four million speakers across Djibouti, Ethiopia, and Eritrea. No published ASR model, no public training corpus, no benchmark. Nothing.

That is on the order of a hundred million people whose first language ranges from underserved to entirely absent in speech technology.

What the gap looks like in practice

The abstraction “low-resource language” obscures what the consequence actually is, so it is worth making concrete.

It means a person who speaks Somali fluently but reads French with difficulty cannot use a voice interface to access a government service, because the voice interface does not understand them.

It means decades of radio and television archives across the region are effectively unsearchable. The audio exists. Finding the moment where a particular person spoke about a particular event requires someone to listen to all of it.

It means a doctor and a patient who do not share a language depend on whoever happens to be available to interpret, with no automated support and no record.

It means every product built on top of speech recognition — captions, transcription, search, voice assistance, accessibility tooling — simply does not exist for these languages, and will not until the layer beneath it does.

Why the market does not fix it

Language technology coverage tracks commercial addressable market fairly closely, and the arithmetic for these languages does not work out.

Building speech recognition for a language requires assembling transcribed audio at scale, which requires access to the audio, native speakers who can transcribe accurately, and money and time to coordinate both. For a large provider evaluating that investment against a speaker population with limited enterprise software spending, it does not clear the threshold. So it does not get built, and the absence is stable rather than temporary.

Waiting for the market to reach these languages means waiting indefinitely.

Why the archives matter

The most useful thing about the Horn of Africa specifically is that the raw material largely exists already.

National and regional broadcasters across Djibouti, Somaliland, Somalia, Ethiopia, and Eritrea have produced radio and television programming in local languages for decades — in some cases since before independence. News, interviews, cultural programming, music, public affairs, oral history. Thousands of hours accumulated year on year.

What that material lacks is transcription. It is catalogued by programme and date, not by word, which makes it inaccessible to search and useless as training data in its current form.

The work required is not conceptually hard. It is access to the archives, native-speaker transcription capacity, careful handling of audio that in some cases is decades old, and the sustained effort to do it properly. That combination has simply not been assembled before.

Preservation and capability are the same problem

We think the framing matters here.

There is a version of this work that treats broadcast archives purely as a training resource — raw material to be processed into model weights. There is another version that treats the archive as what it also is: a record of a region’s public life, its music, its arguments, its history as it was described at the time.

These are not in tension. The transcription that makes an archive usable as training data is the same transcription that makes it searchable, citable, and legible to students, researchers, and families. A well-transcribed archive serves both purposes at once.

The order in which they are pursued does matter, though. Building the capability without regard for the material produces a model and an archive left exactly as inaccessible as before. Building it as preservation produces both.

Where we start

We started with Somali because it has the largest speaker population in the region and because a small amount of public data existed to establish a baseline against. That work is published as SASR Eval.

Afar comes next, and it is the harder case — there is nothing to compare against, no dataset, no prior model, no benchmark.

After that, the same approach extends to the other languages of the region. Each one requires its own archive relationships, its own native-speaker transcription capacity, and its own careful evaluation. There is no shortcut that covers all of them at once.

We are a research project based around Djibouti, working outward. If you work with language material from the region — broadcast, academic, cultural, or otherwise — we would like to hear from you.

See our published Somali ASR benchmark, or get in touch about Horn of Africa language material.