Somali has a data problem. Afar has an absence.

Afar is spoken by somewhere between two and four million people across Djibouti, Ethiopia, and Eritrea. It is one of the two national languages of Djibouti alongside Somali. It has a written form in Latin script, a national broadcaster that has produced Afar-language programming for decades, and a living literary and oral tradition.

It has no speech recognition system.

Not a poor one. Not an experimental one. As far as we can determine, there has never been a published automatic speech recognition model for Afar, nor a public dataset suitable for training one, nor a benchmark against which such a model could be measured.

What “no coverage” actually means

When we say a language has no ASR coverage, it is worth being specific about what has been checked.

Whisper’s language list does not include Afar. Neither do the commercial APIs from the major cloud providers. Meta’s Massively Multilingual Speech project, which extended coverage to over a thousand languages, does not include usable Afar ASR. Searching the major open model repositories for Afar speech models returns nothing.

There is no Afar entry in Common Voice. There is no Afar speech corpus in the standard linguistic data archives at a scale useful for training. Academic work on Afar exists in descriptive linguistics — grammars, phonological studies, lexicons — but not in speech technology.

The practical consequence is that if you have an hour of recorded Afar and you want it transcribed, your only option is a human who speaks Afar.

Why no one has built it

The economics are straightforward and unforgiving.

Building an ASR system for a language requires transcribed audio — typically tens to hundreds of hours, produced by people who speak the language and can write it accurately. That is expensive, slow, and requires access to both the audio and the speakers.

For a commercial provider, the return on that investment is calculated against the addressable market. Two to four million speakers, concentrated in three countries with limited enterprise software spending, does not justify the cost of assembling a corpus from scratch. So it never gets built, and the absence persists indefinitely.

This is not a judgement about the companies involved. It is a structural feature of how language technology gets funded. Languages get coverage roughly in proportion to the commercial value of their speaker base, which means a large number of languages with millions of speakers each will simply never be reached by the market.

The archive problem

There is a further difficulty specific to languages like Afar, and it is the one we find most interesting.

The audio exists. It has existed for decades. National and regional broadcasters across the Horn of Africa have been producing Afar-language news, cultural programming, interviews, and music since the mid-twentieth century. Some of that material is preserved on tape, some has been digitised, and much of it sits in archives that were built for broadcast purposes rather than research.

What does not exist is the transcription. Broadcast archives are catalogued by programme, date, and topic — not by word. The audio is there; the paired text that would make it usable as training data is not.

So the raw material for an Afar speech recognition system exists in principle, and has for fifty years, but the work of turning it into a usable corpus has never been done.

That work is not technically difficult. It is a matter of access, native-speaker transcription capacity, and someone deciding it is worth doing.

Why it matters beyond the technology

An absence in speech technology becomes an absence everywhere downstream.

If there is no Afar ASR, there are no Afar captions, no Afar voice interfaces, no searchable Afar archives, and no Afar in any product built on top of speech recognition. Every subsequent layer of language technology — translation, summarisation, search, voice assistance — inherits the gap.

For a population that speaks Afar as a first language and encounters administrative and digital life in French, Amharic, or Arabic, this compounds an existing exclusion rather than relieving it.

There is also a preservation argument. Broadcast archives degrade. Tape deteriorates, formats become unreadable, and institutional memory of what a recording contains disappears with the people who made it. Audio that is never transcribed is audio that becomes progressively harder to find, and eventually is not found at all.

What we are doing

Our first published work is an evaluation of Somali speech recognition. Afar is the harder and, we think, more important case, precisely because there is nothing to compare against.

The first Afar ASR result will be, by definition, the best in the world — because it will be the only one. That is not a boast. It is a description of how empty the field is.

We would rather it were not empty. If you work with Afar language material — broadcast, academic, cultural, or otherwise — we would like to hear from you.

See our published Somali ASR benchmark, and get in touch if you work with Afar or Horn of Africa language material.