Voice agents
Measuring darija transcription: the protocol, and the figure we do not publish
Nobody publishes an error rate for Algerian darija, us included. Here is the protocol that would produce one, and what to know before reading it.
This article was meant to publish measured error rates across several darija transcription models. It contains none, and the reason is the only one worth having: we have not measured them.
What exists instead is more useful than anything we could have invented — the full protocol, with its traps, its normalisation rule and its scoring sheet. A business can run it on its own calls in a day and obtain a figure that concerns it, which no published average does.
It also says why word error rate, as used everywhere, is the wrong unit for this language, and what to replace it with. It is a question of method before it is a question of technology.
Why that figure does not exist
A transcription error rate is computed against a reference: a text established by hand, word for word, by somebody who listened to the recording. Without a reference there is no measurement — there is an impression.
For Algerian darija those public references barely exist. Reference corpora for Arabic cover standard Arabic and a few dialects, and ours is generally not among them; suppliers announce performance "in Arabic" that describes a different language from the one on your calls.
We have not built one either, and that is a priority choice rather than an oversight: building a usable reference requires real recordings, permission to use them, and hours of human transcription. Section 9 covers the legal part, which is the most binding.
The rest of this article is therefore the recipe. We publish it whole, including the places where we know it is expensive, because an incomplete protocol produces figures that compare badly — which is worse than no figure at all.
What "error rate" means
The standard measure is word error rate. Between the automatic transcription and the human reference you count substituted words, deleted words and inserted words, and divide that total by the number of words in the reference.
It has two properties worth holding in mind. It can exceed a hundred percent, because a talkative system inserts more words than were there. And it treats all errors as equivalent: confusing two function words and getting an order number wrong count the same.
It is that second property that makes it misleading for professional use. A system that transcribes everything correctly except names and references will have a flattering rate and be unusable, because the words it misses are exactly the ones that trigger an action.
From that follows a way of reading rates published by others, where they exist. The first question is not "how much" but "on what": which sample, which language exactly, which reference, which date. A rate missing those four answers does not compare to yours, and quoting it means borrowing somebody’s conclusion without their experience.
The rate remains useful for comparing two systems on the same audio. It does not say whether a system is good enough for your use — those are two different questions, and the second is measured differently.
The word is not the right unit here
Written darija has no standardised spelling. The same word is written several ways depending on who transcribes, and a word-level measure counts those differences as errors although no information is lost.
The consequence is that a rate computed without precaution mainly measures the spelling disagreement between your human transcriber and the model. Two people transcribing the same call already produce different texts, and the gap between them can be of the same order as the gap you are trying to measure.
The remedy is a normalisation rule written before starting and applied on both sides: both texts go through the same transformation — the same handling of short vowels, spelling variants, punctuation, numbers written out. It is not a correction, it is a levelling.
We also recommend a second measure, and it is the one that decides: the error rate on **decisive elements**. Call by call, you list the pieces of information that trigger something — a name, a number, a quantity, a date, an address — and count how many are transcribed correctly. A smaller measure, slower to establish, and much closer to the real question.
The sample: how many, and which
A usable sample is made of real calls, taken as they come, over a continuous period. Recordings chosen because they are clear produce a measurement describing a world where everybody speaks well.
The size is reasoned in minutes of audio rather than in number of calls. Around thirty complete conversations gives enough material to separate two very different systems; below that, the measured gap and chance look too alike to conclude anything.
Composition matters more than size, and it is the most neglected point. A sample has to contain short and long calls, good and bad lines, men’s and women’s voices, people who speak fast, and at least a few calls made from somewhere noisy — a market, a car, a workshop.
Note, finally, the calls you exclude and why. A sample whose exclusions are not written down is not reproducible, and it is the first thing somebody will ask for if they want to redo your measurement.
The reference: who transcribes, and how
The human reference is the expensive part and it decides the value of everything else. It requires somebody whose darija is your customers’ — not merely somebody who speaks Arabic.
The instruction given to the transcriber has to be written and short: transcribe what was said and not what should have been, keep repetitions and hesitations, correct nothing, mark inaudible passages rather than guessing them. A transcriber who improves the text builds a reference even a perfect system cannot reach.
The pace deserves announcing so that the budget surprises nobody: transcribing spontaneous speech takes several times the duration of the audio, more still when the line is bad and passages have to be replayed. Thirty conversations are not transcribed in a morning, and a team promised otherwise produces a rushed reference, which ruins the measurement more surely than a bad model does.
Double entry is the precaution worth its cost on part of the sample: two people transcribe the same five calls independently, and their disagreement is measured. That figure is the floor of your measurement — no system can do better than the agreement between two humans on the same audio.
That floor is also what makes a comparison honest. A gap between two models smaller than the disagreement between your two transcribers proves nothing, and it is the result most published comparisons never report.
What to record around the audio
Every call in the sample carries a record card, and it is filled in at collection time because it is impossible to reconstruct afterwards. Four fields are enough most of the time.
The duration, in seconds. The line quality, on a three-value scale judged by ear — good, medium, bad. The presence of background noise, yes or no. And the channel: landline, mobile, messaging app, each with its own compression.
Those fields serve one thing and it is decisive: they let you tell whether a system is bad in general or bad on bad lines. Those are two different problems, one solved by changing supplier and the other by changing operator or equipment.
Add to them, if you can, the speaker’s region as you know it. Regional variation is real, and a system can be excellent on one city and mediocre on another without the average showing it.
The scoring protocol, step by step
First, freeze the sample: the audio files and the record cards, in a folder nobody touches again. Any modification after the start invalidates comparisons made before it.
Second, produce the human reference, then apply the normalisation rule to that reference and set it aside. Third, run the same audio through each system, on the same day, and normalise each output with exactly the same rule.
Fourth, compute word error rate for each system, then count the decisive elements separately. Fifth, write the two sets of results side by side with, at the head of the table, the date, each system’s version and the human disagreement measured at the double-entry stage.
Sixth — and this is the step everybody skips — run one of the systems a second time, the next day, on the same files. If the result moves, your measurement has a variability of its own, and it has to be known before concluding that one system is two points better than another.
Comparing without mistaking the comparison
A comparison only means something under identical conditions, and the list of what has to be identical is longer than people think: the same audio, the same normalisation, the same period, and the same absence of a custom dictionary.
That last point deserves insisting on. A system that has been given the list of your products and neighbourhoods is no longer comparable to a raw one, and the measured difference is the list’s, not the model’s. That reconciliation layer is often the right solution — but then it is measured separately, as an improvement applied to the best system.
The honest result of a comparison is therefore rarely "system A is better". It is rather: on this kind of call, in these conditions, on this date, A loses fewer decisive elements than B, and the gap exceeds the disagreement between our transcribers.
There is finally one request to put to the supplier, and only one is worth making: the exact version of the system used, with its date. Many will answer that they do not disclose it, which is already information — it means your measurement will not be reproducible at their end, and that degradation from one day to the next can never be held against them.
Write the date out in the table. Transcription models change without notice and without changing name, and an undated table describes a state of the world that no longer exists.
What we have observed, and what we have not
Here is exactly what we hold, so that nobody credits this article with more than it contains. Our own transcription chain runs with three evaluation cases, written to check that the service answers and returns plausible text.
That is not a measurement. Three cases are not a sample, there is no human reference behind them, and nothing has been compared against a second system. We therefore have no error rate to publish, in darija or in any other language.
What we do have is qualitative, repeated and dated: across the work carried out to date, what breaks is not the language but the lexicon — proper names, neighbourhoods, product references and numbers. It is a field observation, described in full elsewhere, and it does not turn into a percentage without the protocol above.
We will publish a figure the day we have run this protocol on a real sample, with the agreement of the people recorded, and we will publish the method and the table together. Until then, the weekly listening session remains the most honest way to know where you stand.
Before recording anything
A customer call is personal data, and recording it is processing. Law 18-07 of 10 June 2018, amended and completed by law 25-11 of 24 July 2025, applies to this sample in full.
Three concrete obligations follow before the first minute is recorded. Inform the person at the start of the call, in a sentence they understand and in their language. Limit retention to the duration of the measurement and write that down. And enter the processing in your register, with its real purpose — evaluating the quality of a transcription system — and not a general one.
The part businesses forget is what becomes of the sample after the measurement. A folder of real recordings sitting on a workstation is the kind of thing that makes an inspection very short and very bad. Decide the deletion date before collecting, not after.
If the audio has to leave the territory to be transcribed — which is the case for most services — the question changes nature and becomes one of data transfer. It is the point in the protocol that stops the most projects, and it is better that it stops them before collection.
The check: the short version
The full protocol takes a day’s work and a legal agreement. There is a one-hour version that gives no error rate but answers the question "is it worth going further".
Take twenty recent calls. Listen to them, noting only, for each, the decisive elements spoken — the name, the number, the quantity. Have the same twenty transcribed by the system you are considering. Count how many of those elements are correct.
You get no measurement comparable to anything, and that is not the point. You get the one piece of information that decides at this stage: the proportion of the information that triggers an action and arrives intact.
If it is bad, no tuning will save it and the full protocol is a wasted expense. If it is good, you know the serious measurement is worth the day it costs, and you already know which calls to put into it.
What we do, and what we refuse
We run this protocol with you, from building the sample to the final table, and we leave you the files, the normalisation rule and the scoring sheet — so that you can redo it without us in six months and compare.
We refuse to publish an error rate we have not measured, including rounded, including "in our experience". A percentage on this page would be quoted for years by people with no way of knowing where it came from, and that is exactly what this protocol exists to prevent.
We also refuse to measure on recordings collected without informing the people concerned. That is not an abstract moral position: it is unlawful processing, and a figure obtained that way can neither be published nor shown to a client.
What you can do without us is section 11’s short version, and the best outcome would be for you to run the full protocol and publish it. This market lacks an honest measurement, and the first one to appear with its complete method will serve everybody, us included.
Frequently asked questions
Why not use the figures suppliers publish?
Because they almost always concern standard Arabic or a set of dialects that does not include ours, and because the protocol is not published with them. A percentage with no sample, no reference and no date compares to nothing.
How many calls are really needed?
Around thirty complete conversations is enough to separate two clearly different systems. What is most often missing is not the number but the variety: bad lines, background noise, fast speech, several regions.
Can old recordings be used?
Technically yes, legally only if they were collected for a purpose that covers this use. Reusing calls recorded for staff training in order to test a supplier is a change of purpose, and it is handled before, not after.
Should we transcribe ourselves or bring somebody in?
It does not matter who, provided their darija is your customers’ and they follow a written instruction. The main risk is not competence, it is the temptation to correct what was said — an improved reference makes the measurement uninterpretable.
Can the error rate really exceed 100%?
Yes, and it is normal: insertions count, and a system that hallucinates text over silence produces plenty. It is even a useful signal — a rate above a hundred indicates filling behaviour rather than a simple lack of accuracy.
Will you publish your results?
Yes, the day they exist, with the full method, the sample described and the date. Until then this page contains no figure, and its absence is deliberate rather than an oversight.
Where we come in
The twenty-call short version says within an hour whether the serious measurement is worth the day it costs.
- We build the sample with you, exclusions written down, and cards filled in at collection.
- The normalisation rule is fixed before the first transcription and applied on both sides.
- The files, the scoring sheet and the dated table stay with you, redoable without us.
No percentage will come from here until we have measured it ourselves on a real and declared sample.
Read next
The cost of a call, minute by minute
Four meters run during a voice call and none of them counts the same thing. Confusing them gives an estimate wrong by a factor, not by a margin.The words transcription does not know
The voice agent understands the sentence and gets the name wrong. That is the opposite of what people fear, and it is what makes an appointment unusable.Voice agent: the thirty-second test
A voice agent does not replace your switchboard. It takes the calls nobody takes — and only the ones that fit in thirty seconds.
Let us talk about your project
A free audit, no commitment: we look at your online presence and tell you what is holding it back.