Skip to content
Client login

Free Audit

Artificial intelligence

A chatbot that understands darija: what it adds to the work

Darija is not expensive because it is difficult. It is expensive because it is never written the same way twice, and that is settled on your own messages.

Published on 26 August 2026 — Algeria Agency

A shopkeeper in Bab El Oued receives about a hundred messages a day and wants to answer faster. He asks for an assistant "that understands darija", because that is what his customers write. The first question put to him is not which of the three languages he wants, it is to show us a hundred real messages.

That is where the subject becomes concrete. A modern artificial intelligence system copes with a great deal of darija; what it does not do is guess how your customers write the words that matter to you. The extra work is there, and it is measured in days of tuning rather than in a feature to buy.

This article describes what darija adds to the work, what it breaks when added badly, and the three cases where French alone is enough. It does not repeat what an assistant must refuse to answer, which is the prerequisite, and it does not discuss price: the reason is in section 5.

Three capabilities sold under one word

"Understanding darija" covers three different things, and the first decision is saying which one you are buying. The first is understanding darija written in Latin characters, with digits standing in for sounds the alphabet does not carry. The second is understanding darija written in Arabic characters. The third is replying in darija.

They share neither a cost nor a risk. Understanding both scripts is a recognition problem: the system has to tie several spellings of one word to a single intent. Replying is a brand decision, handled in section 6, and many businesses that need the first have no interest in the third.

A supplier who does not make you choose is selling all three and usually delivers one. The question to ask in a demonstration is narrow: show me the system understand the same message written three ways, then show me what it replies, and tell me which of the three is adjustable after go-live.

The rest of this article is about the first, because it is the one with a measurable effect on your daily work and the one that takes the most tuning.

Your messages are the only raw material

Nobody can cost this subject without reading your messages, and a quote given before that reading describes some other business. An industrial bakery, a recruitment office and a spare-parts dealer receive three different darijas, because the vocabulary that causes trouble is yours, not the language.

What has to be gathered is one operation: export or copy a hundred real inbound messages, as written, without correcting them. The mistakes, the abbreviations, the digits inside words and the transcribed voice notes are part of the material; cleaning them up is revising for an exam that will not be set.

Those hundred messages serve three times. They say what share of your traffic genuinely needs darija — section 11 turns that into a check. They supply the list of variants that matters. And they become the test set every later change is verified against, which is the only protection against an improvement that breaks something else.

They are also the first sort between suppliers. One who asks for your messages before talking about a solution has understood the problem; one who talks about a model before reading a line is talking about something other than your business.

There is no spelling, and that is where the work is

Written darija has no standard, and that is not a defect: it is a spoken language its speakers write by ear. So the same word arrives in several Latin spellings, several Arabic ones, with or without digits, with vowels written or not, split or joined.

For a system the consequence is not that it fails to understand: it is that it fails to match. Two customers asking for the same thing produce two different forms, the system recognises one and not the other, and the log of unanswered questions fills with one request written six ways.

The work is therefore to build a list of variants for the words that decide things in your business — products, commercial gestures, the words for availability and price — and then to set a matching threshold. It is not a translation list and must not be mistaken for one: it ties spellings to an intent, not a word to a word.

The surprising part is that the list cannot be guessed. Three people who speak darija perfectly will produce three different lists, and none will look like what your customers actually write. That is why the previous section comes before this one.

What a generic model already does

Part of the work you will be billed for no longer needs doing, and it is honest to say so. Recent models understand an ordinary darija sentence, the structure of a request, a negation, a question about availability, including written in Latin characters with digits.

What they do not do is narrower and more awkward: they fail on the names that belong to you — references, local brands, neighbourhoods, in-house terms — and they fail on the words whose spelling varies most, which are often the most frequent in your trade.

The practical consequence is an order of work. You do not start by "training a model on darija"; you start with the layer that matches what the model understood against what your catalogue contains. It is the same mechanism described in the words transcription does not know, applied to writing.

It also moves the supplier question. A larger model rarely improves those cases; a maintained list improves all of them. A quote offering one without the other has chosen the most expensive line and the least effective.

The cost is a tuning cost, and there is no price in this article

What darija adds is counted in days of tuning and in review cycles, not in a licence. Gathering and reading a hundred messages, writing the first list of variants, setting the threshold, testing the result against the hundred, correcting, then doing it again after two weeks of real use: that is the object of the extra work.

We give no dinar range here, and the reason deserves writing down rather than guessing at. A range would be about our own working days, which is a rate card, and this company publishes none; a figure carried over from a foreign supplier would describe another market and another job. What we publish instead is the line items, so you can ask for them one by one from whoever you consult.

The useful thing about the amount is its shape rather than its level: darija is a start-up cost and a maintenance cost, never a single one. A quote carrying only an installation line describes a system that will degrade quietly, because your customers’ vocabulary moves.

So ask every supplier for three separate lines: the initial tuning, the revision after some weeks of use, and the yearly upkeep. A quote that refuses to separate them cannot be compared with another, which is often the intention.

Replying in darija is a brand decision

Understanding darija and replying in it are two separate decisions, and the second is not technical. A reply written in Latin darija reads as close and familiar to some of your customers and as careless to others — the same sentence, read by an institutional client or a supplier, says the business writes the way it speaks.

The position we advise is asymmetric and easy to hold: understand all three languages, reply in French or in standard Arabic according to the message’s dominant language, in a short and concrete register. The customer does not need to read you in darija; they need to be understood when they write it.

There is a clear exception, and it is the transcribed voice note or the very informal exchange on a messaging app where your brand is already familiar. There, replying in the customer’s register is coherent — and the channel already imposes constraints of its own before that one. It is a decision taken once, written down, and applied everywhere — not a per-channel setting.

What must never happen is the system choosing its own register message by message. The variation is visible to a customer who writes twice, and it gives the impression of talking to two different businesses.

The three cases where French alone is enough

The first is business-to-business trade. Exchanges with a buyer, a supplier or an administration happen in French or standard Arabic, with stable vocabulary and longer messages. Adding darija to that flow is paying for a capability nobody will use.

The second is the short message with a reference. If your customers write a product reference, a size, an order number and three words around them, what decides the reply is the reference, not the language. You have a catalogue-matching problem, and darija will not change it.

The third is less expected: when your team already replies within minutes. An assistant’s benefit comes from the time it gives back, and if your messages are handled as they arrive, darija only adds one more intermediary between a customer and a person who would have understood them perfectly well. The check in section 11 will tell you which of the three you are in.

Those three cover a serious share of the businesses that ask us about this. Saying so is more useful than a demonstration: the bad purchase here is not a system that fails, it is a system that works and serves no purpose.

What breaks when darija is added

The addition is not neutral for everything else, and that is the part demonstrations never show. Widening recognition to accept several spellings of a word makes the system more tolerant, and therefore more inclined to match two requests that have nothing to do with each other. In French that shows up as off-topic replies to messages that worked perfectly the day before.

The setting responsible is the matching threshold, and it has no universally good value. Too strict and darija is not recognised; too loose and French messages go astray. It is a trade-off to be measured on your hundred messages, not chosen from a configuration box.

There is a second, slower effect: the list of variants grows, and a list that grows becomes a list nobody rereads. Entries added in a hurry for one case stay, and six months later they produce matches nobody can explain.

The discipline that answers both fits in one sentence: every change is replayed against the hundred messages before go-live, and the result is compared with the previous one. Without a test set, an improvement and a regression look exactly alike.

Upkeep: a living list, and somebody who holds it

Commercial vocabulary moves: a product launches, a promotion creates a phrase, an expression spreads, a neighbourhood acquires a nickname. A frozen list of variants loses reach every season without anything saying so — the log fills a little more, and nobody makes the connection.

The useful rhythm is quarterly and the work is short: reread the unanswered questions since last time, take the new forms from them, add them, replay the test set. An hour or two, provided the log exists and somebody owns it.

That ownership has to be named inside the business, and it is the part most readily and most badly delegated. A supplier can hold the mechanism; they will not know that a new expression has come to mean your product among your customers, because they are your customers.

A business that cannot name anybody for that quarterly hour should stay with the three cases in section 7 and handle its messages by hand. That is an acceptable answer and it beats a system that degrades.

Understanding in writing and understanding by ear are not one subject

The confusion is common and expensive, because the two capabilities are sold together. In writing, the system receives exactly what the customer typed: the errors are errors of interpretation. By ear, it receives what a transcription believed it heard, and the errors were committed before any language model was involved.

The practical consequence is an order: if your customers mostly send voice notes, your problem is not the chatbot, it is transcription — and transcription decides before the model, which is covered elsewhere.

There is a mixed case, very common here, and it deserves naming: the voice note transcribed automatically by the customer’s own app and then sent as text. You then receive text that has already been through a transcription, with its own errors, and the matching has to be more tolerant than for typed text.

The sorting is done by looking at your hundred messages: how many are typed, how many are voice, how many are transcriptions. Those three proportions decide half the project, and they are counted in an afternoon.

The check: a hundred messages, three columns

Take the last hundred inbound messages, across all channels, and fill in three columns: the message’s dominant language — French, Arabic, darija — whether it carries a product reference, and how long your team took to reply.

Two numbers decide. The first is the share of darija messages whose answer does not depend on a reference: that is exactly the population work on darija would improve. The second is your median reply time. If the first share is small, or the delay is already minutes, you are in one of the three cases in section 7 and there is nothing to buy.

Add one reading that costs ten minutes and often surprises: in the darija messages, count how many different words denote the same thing. That is the size of your variant list, and therefore of the tuning. A list of twenty forms is not the same job as a list of two hundred.

This check is also what you hand to a supplier. It turns a conversation about language into a conversation about your traffic, and it makes two quotes comparable, which no demonstration does.

What we do, and what we refuse

We read your hundred messages with you, take the variant list and the test set from them, set the matching threshold on your data, and come back after two weeks of real use. We leave the list and the test set with you, in a format somebody else can pick up.

We refuse to announce a comprehension rate. We have no annotated corpus and no comparison between models, and a percentage offered without them would measure something other than what it claims. The only honest figure is the one your test set produces on your messages, and it belongs to you.

We also refuse to deliver an automatic reply in darija on a channel where your brand speaks otherwise, even on request: the shift in register shows within two messages and costs more than the closeness gains.

What you can do without us is the first half of the work: the three columns in section 11, on a hundred messages, in one afternoon. Many businesses discover there that their difficulty is the catalogue rather than the language, and what the list of what an assistant must refuse says will help them more than darija.

Frequently asked questions

Do we need a model specially trained on darija?

Rarely, and it is the most expensive line in a quote. Recent models handle an ordinary sentence; what fails is the vocabulary that belongs to you, and that is corrected by a matching layer against your own lists rather than by a different model.

Darija in Arabic characters or Latin ones?

Both arrive, often from the same customer depending on device and habit. So both scripts have to be handled, and that is precisely why the work is about variants rather than about a language: one intent exists in two alphabets and several spellings.

How long before it is usable?

The first tuning is counted in days once the messages are gathered. What takes time is the revision after real use: it takes two to four weeks of traffic to see what the first list missed, and that revision is part of the work, not after-sales service.

Can we start in French and add darija later?

Yes, and that is often the healthiest order. The log of unanswered questions accumulated over the first months is exactly the raw material the variant list needs, and you will know by then whether the volume justifies the work.

Our customers mix three languages in one sentence. Is that a different problem?

It is the same mechanism and it is already covered in the article on what an assistant must refuse to answer. What is added here is the spelling variance inside darija itself, which is what makes the matching fail even when the sentence was understood.

What about voice notes?

Count them first. If they dominate your traffic, your subject is transcription rather than the chatbot, and the order of work changes completely. If they are a minority but present, expect their transcriptions to arrive with errors of their own and the matching to need more tolerance for them.

Where we come in

Across your hundred lines, what settles it is the share of darija whose answer does not depend on a product reference: below that, the subject is not language.

  • We read your real messages with you and take from them the list of forms that matters, not a translation list.
  • The reference trial is handed over as it is: any supplier can replay it without coming through us.
  • The tolerance threshold is set in front of you, and the verification is replayed after some weeks of real traffic.

No comprehension rate will come out of here: lacking an annotated corpus, such a percentage would speak of a reality that is not yours, and your trial describes it better.

Read next

Let us talk about your project

A free audit, no commitment: we look at your online presence and tell you what is holding it back.

We measure how this site is used with Google Analytics, to learn which pages actually help. You can stop that measurement at any time from the footer. Cookie policy