Artificial intelligence
Chatbot: what it must refuse to answer
An assistant is judged on its refusals, not its answers. What the list has to contain, and what happens when nobody wrote it.
Every page selling a conversational assistant describes what it can answer. That is the easy part, and it is also the part that decides nothing: a recent model answers almost anything, confidently, including when it is wrong.
What decides whether the project survives is the opposite: the list of what the assistant is not allowed to handle alone. That list is short, it can be written in one meeting, and writing it is the only part of the work that genuinely commits anybody.
This article describes what it has to contain, what happens when it does not exist, and the three or four checks you can run yourself before speaking to any supplier — us included.
Refusing is the feature
An assistant that says "I cannot find that, I am passing you to a colleague" has just done its job. That is counter-intuitive because it looks like a failure, and because whoever sold it to you spent the demo showing the opposite.
The reason is arithmetic. A model that is right nine times in ten is wrong once in ten, and the question is never "how often" but "about what". Being wrong about opening hours costs a correction; being wrong about a delivery date quoted to a customer costs the customer, and sometimes more.
So a system that cannot refuse does not make ten percent errors spread evenly: it makes ten percent errors of which some land on the subjects where a mistake cannot be undone. Refusing is the mechanism that removes those subjects from the draw.
A practical consequence, also counter-intuitive: in month one, a high refusal rate is a good sign. It means the limit is set tight and the system is not taking initiative. You loosen it afterwards, question by question, guided by what the log shows. The reverse — start broad and tighten after an incident — is paid for with a real customer, and the tightening always arrives too late for that one.
The five questions
Before deciding anything, export last month’s private messages and classify them. The exercise takes an hour and almost always produces the same result: five questions cover the large majority of the volume, and they are the five you already suspect — prices, hours, address, lead times, delivery to a given wilaya.
That classification is the real scope. An assistant that handles those five cleanly, and nothing else, captures most of the benefit available. An assistant that tries to cover the long tail of rare questions spends the budget where the volume is not and gets things wrong where the stakes are highest.
The classification also tells you how many of your messages arrive outside working hours. That figure is the only honest commercial argument for this kind of project, and it is yours: nobody else can give it to you, and a supplier who offers you one before measuring it is offering you an average taken somewhere else.
The classification is done by hand, and deliberately so. Automatic categorisation groups by vocabulary similarity and will hand you ‘delivery’ and ‘lead times’ separately because the words differ, when they are the same question to your customer. Somebody who knows the trade does that grouping while reading, and what they learn doing it is often worth more than the final table.
What happens when it invents
A general-purpose model asked for your returns policy will produce one. It will be phrased exactly like a real returns policy, with a plausible window and a professional tone, and it will be wrong. That is normal behaviour for these systems: they are built to produce the most likely continuation, not to check that it exists.
The damage is not the error, it is the credibility. A customer whose message your contact page never answers knows they have no answer. A customer told by your assistant that they have fourteen days to change their mind believes they have fourteen, acts accordingly, and discovers otherwise at the moment it costs you most.
The defence is structural, not rhetorical. You do not obtain caution by asking the model to be cautious: you obtain it by grounding it in a closed corpus — your pages, your catalogue, your procedures — and making it return the source of every answer. An answer without a source does not ship.
One point suppliers happily skirt is worth insisting on: you do not get this behaviour by instructing the model to be careful. A directive along the lines of ‘only answer if you are certain’ improves the statistics without removing the case, because the model has no internal measure of its own certainty to consult. What removes the case is cutting the access: no source, no answer, enforced in code.
Messaging is the counter
The internet market observatory published by ARPCE counts, for the second quarter of 2025, some 59.10 million internet subscriptions in Algeria, of which 88.71% are mobile and 11.29% fixed.
That proportion explains why an assistant confined to the website misses the subject. Access that is overwhelmingly mobile is access where the messaging app is permanently open and the browser is open intermittently, and where writing a private message takes two gestures while filling in a contact form takes ten.
In practice: the commercial conversation happens in WhatsApp, Messenger and Instagram direct messages, and it happens there whether your site is good or bad. An assistant project that does not cover those channels automates the least-visited part of your presence.
ARPCE, internet market observatory, second quarter of 2025
Three languages in one sentence
Your customers do not write in one language, they write in three, often inside a single sentence. A typical enquiry mixes darja transcribed in Latin script, a French word for the product, and numbers spoken in Arabic or French depending on the person’s habit.
This is where most demos lie by omission. They are run in clean French, on complete and well-punctuated sentences, and they predict nothing about how the system behaves on your real conversations. The quality gap between those two regimes is the first thing to measure, before price.
The measurement is simple and you can insist on it: replay a hundred of your old messages, compare the assistant’s answers with the ones your team actually gave, and count the differences. A supplier who refuses that test, or offers to run it after go-live, is telling you something useful.
One particular case is worth flagging because it turns up everywhere: brand names written as they sound. The same product will be searched for in three or four spellings, none of them the one on your listing, and an assistant that does not map them will reply that it does not know an item sitting on the shelf. Those spellings can be pulled out of your own conversations in an hour; they cannot be guessed. That survey is the start of a job in its own right, and what darija really adds is counted from it.
The list written before go-live
The list fits on one page and always contains the same families: anything that commits a price, anything that commits a date, anything touching a complaint or a dispute, and anything amounting to regulated advice — medical, legal, financial.
Beside each line there has to be a name, not a department. "Passed to sales" is not a procedure; "passed to whoever holds the delivery calendar, and failing that to the shop manager" is one. Without a name, the handover lands in a shared inbox that everybody treats as somebody else’s.
That page has to be signed before go-live, not discovered after. It is also the document you will re-read in six months to understand why the system does what it does, and the only deliverable of the project that stays legible to somebody who was not there.
The list gets re-read, and the right rhythm is quarterly. It ages in one specific, predictable direction: businesses loosen over time, because each handover visibly costs somebody time while the risk avoided is never seen. A dated review, with the person who picks up the transfers, is what stops the limit dissolving over six months without any decision having been taken.
The handover that does not restart the conversation
A bad handover cancels the benefit of the assistant. The customer has explained the problem, the assistant has reached its limit, and the person picking up opens with "hello, how can I help". From the customer’s side, they have now wasted their time twice.
A proper handover carries three things: the whole conversation, the reason for the transfer, and whatever the assistant had already established — the product reference, the order number, the wilaya. That is a requirement to set at scoping, because it is easy to build at the start and expensive to retrofit.
There is an interaction with opening hours best settled at the outset. A handover triggered at eleven at night reaches nobody, and if the assistant does not say so the customer waits for a reply that will come next morning without their knowing it. The rule that works is simple: outside hours, the assistant announces the transfer and the time it will be dealt with, rather than presenting it as immediate.
The log of unanswered questions
Every time the assistant refuses, the exchange should be recorded with the question that was asked. After a month, that list is the most useful deliverable of the project, and it has nothing to do with artificial intelligence.
What it contains is your missing pages. The questions your customers genuinely ask and that nothing on your site answers, ranked by frequency, in the exact wording they use. Marketing teams pay well to obtain that list by other means, and you get it as a by-product.
It has a second, less pleasant use: it measures the system’s coverage honestly. If the log does not empty out over the months, the assistant is not improving, and that needs saying rather than recalibrating the indicators around the result obtained.
The concrete use fits into a twenty-minute monthly meeting. Take the ten most frequent questions in the log, decide for each whether it deserves a page, a sentence added to an existing page, or nothing at all, and note who writes it. Three months of that discipline usually empties half the log, and what remains is what the assistant will have to keep transferring.
What it costs, and why nobody tells you
These systems bill by usage, to foreign suppliers, in currencies you do not earn. The cost depends on the number of conversations, their length, and how much of your content is consulted for each answer — three variables, none of which is known before go-live.
We deliberately put no chart in this section. A per-request price published today is wrong in six months: suppliers revise their tables without notice, withdraw models from the catalogue, and bill in euros or dollars, so a figure dated to one quarter does not travel to the next. A curve would give that instability the appearance of a trend.
What you can insist on instead is a ceiling. A monthly threshold set in advance, an alert before it is reached, and a decided behaviour beyond it — falling back to message-taking rather than an invoice that keeps running. That is a contract clause, not a technical feature, and it is worth more than any estimate. The ceiling bounds usage; what moves the go-live price sits elsewhere, and five scope decisions account for it.
Opening hours, said honestly
An assistant available at night creates an expectation that has to be defused explicitly. If it answers at eleven on a Friday evening without saying when a person will take over, the customer assumes somebody is watching, and the absence of a human reply next morning is experienced as abandonment rather than as closing time.
The wording costs one sentence: what the assistant can do now, what will wait for opening, and when that is. Businesses that remove it to look more available get the opposite of what they want, because an implicit promise not kept is remembered better than a stated limit.
The local detail that matters is Friday and public holidays. An assistant announcing ‘we open tomorrow at eight’ on a Thursday evening is a full day wrong, and that kind of error is received as inattention rather than as a bug. The closure calendar has to be data somebody at your end updates, not a rule written into the system once and forgotten.
When a page is enough
If your five main questions are stable, their answers do not change month to month, and nothing on your site addresses them, then your problem is not an assistant problem. It is a missing page, it costs a day to write, and it fixes private messages and Google searches at the same time.
This case is common. A business fielding the same lead-time question twenty times a week does not need a system capable of understanding natural language; it needs to write its lead times down somewhere and point at them. The assistant becomes useful above a certain volume and variety, not below it.
The check takes two minutes: take your five questions and look for their answers on your own site, the way a customer would. If you cannot find each one in thirty seconds, start there. A supplier who does not make you run that exercise has an interest in you not running it.
What we do, and the figure we will not give
We read your past conversations, we write the refusal list with you, we build the assistant on your content, and we watch it for two weeks before letting it run. We re-read the refusal log with you monthly. That is the whole offer.
The figure we will not give is the containment rate — the share of conversations that end without a human. We have it for the systems we run, it is real, and publishing it as an average would be dishonest: that rate is decided by the mix of questions each client receives, not by the quality of the build. A business whose customers mostly ask about opening hours will get a high rate from any supplier; a business whose customers negotiate prices will get a low one from the best. An average between the two measures our clients rather than our work, and you could do nothing with it.
What we will do instead is measure it at your end, on your own messages, before you have signed anything. And there is one case where we will tell you not to do the project: the one in the previous section. If your five questions fit on a page you have not written, write it — we do not bill for that conversation, and we would rather lose it cleanly than sell a system to compensate for a missing text.
Frequently asked questions
How long before an assistant is usable?
A few weeks, half of it spent reading your conversations and writing the refusal list. The technical part is the shortest stretch of the project, which often surprises people.
Can it take a complete order?
It can prepare one and present it for human confirmation. It does not confirm alone: a wrong quantity or address only surfaces at delivery, when it costs the most.
Does it really understand darja?
Partly, and very unevenly by subject. It is measurable on your own messages before any commitment, and it is the first measurement to insist on.
What happens if the AI supplier becomes unavailable?
The assistant should fall back to taking a message and say so. That behaviour is decided at scoping; with no decision, the default is an error page in front of a customer.
Are our conversations used to train a model?
That depends on the supplier and the contract, and it is a question to ask explicitly. The answer should be written into your contract, not inferred from a help page.
Can it be switched off without losing anything?
Yes, if the indexed content is yours and stays exportable. That is a condition to check beforehand, because afterwards it is no longer negotiable.
Where we come in
A month of messages sorted by question, discarding nothing, is the only material worth having. That sorting does not yet say what can be answered without you.
- We go through your earlier exchanges before drafting a single reply.
- We draw up with you the list of what it must decline, which runs long.
- We measure at your company, on your messages, before you commit to anything.
We will give you no autonomous-handling rate from anywhere else: that figure depends entirely on your own questions.
Read next
Two quotes for the same chatbot: what actually moves the price
Two suppliers price the same need at one and at five. The gap almost never comes from technology: it comes from five scope decisions.A chatbot that understands darija: what it adds to the work
Darija is not expensive because it is difficult. It is expensive because it is never written the same way twice, and that is settled on your own messages.Which messaging channel, and what each one imposes
The channel decides the rules, the deadlines and what you are allowed to send. What to know before choosing where to reply.
Let us talk about your project
A free audit, no commitment: we look at your online presence and tell you what is holding it back.