Skip to content
Client login

Free Audit

Digital marketing

A/B testing: what your traffic lets you conclude

The question is not which version wins, but whether you have the volume to know. Nearly always the answer is no — and there is better work available.

Published on 13 June 2026 — Algeria Agency

The article on conversion rate sets the limit and the one on landing pages points at it: comparing two versions takes a volume most Algerian sites do not have. This page starts from there.

A test is not an opinion settled by numbers. It is a procedure answering one narrow question — is this difference real or chance — and it answers nothing else. Used outside its domain it produces random decisions presented as a method.

The domain is defined by volume. Below it you do not get a weak result: you get something that is not a result, and experience shows somebody always reads it as a win.

This article says how to know which side of the limit you are on, what to do on each side, and why we will not publish the figure everybody asks for — because that figure does not exist independently of you.

What a test can answer, and nothing else

A test compares two versions seen by two groups of visitors drawn at random over the same period, and it answers one question: is the observed gap large enough not to be explained by chance?

It does not say why. A winning version does not teach you what produced the gain, especially if you changed four things at once — which is the case in nearly every test we see presented.

Nor does it say whether the winner will stay the winner. A result obtained during a promotion, a school term or a particular month describes that period, and nothing guarantees it holds next quarter.

And it says nothing about value. A version producing more orders and more refusals at delivery loses money while winning in the tool, because the tool stops at the place the conversion-rate article describes as the wrong one.

Those four limits do not make the method useless. They define a precise question it answers well, provided the volume is there — which is the next section.

The number we will not publish

The question asked in every first meeting is "how many visitors do we need?". There is no general answer, and the absence of one is not evasion: it is the nature of the calculation.

The volume needed depends on two things that belong to you. Your current rate — the lower it is, the more people it takes to distinguish a gap. And the size of the effect you want to be able to detect — a doubling shows up fast, a ten per cent improvement takes an order of magnitude more.

That is why there will be no chart on this page. A single figure published as the answer would be wrong for every reader at once, each in a different way, and it would be quoted for years.

What can be said without error fits in one sentence of common sense: if you are not counting your conversions in hundreds within each version, you will not be able to distinguish a modest improvement from noise. And modest improvements are most of what anybody is looking for.

The practical consequence is a test before the test: take last month’s conversion count, divide it by two, and ask whether that number looks solid to you. If it fits in one hand, the answer is settled.

Looking too early: the error that produces winners

The most destructive fault is not lack of volume, it is daily checking. A test looked at every day always ends up showing a gap, because chance produces gaps constantly.

The mechanism is easy to feel: if you toss a coin a hundred times and stop the moment heads is six ahead, you have manufactured proof that the coin is loaded. That is exactly what a dashboard checked every morning does.

The rule is therefore to fix the duration and the volume before launching, to decide nothing before, and not to stop because one version has pulled ahead. Discipline matters more here than anywhere else in this pillar.

The second rule of the same kind is not to relaunch a lost test hoping for better. Repeating until the desired result appears is the same error, spread over weeks and much harder to spot in a report.

One unpleasant consequence for an agency: most honest tests end with no winner. A report that always announces one describes a reading practice, not a series of discoveries.

Whole weeks, never days

Buying behaviour is not the same from Sunday to Thursday as at the weekend, nor in the morning as in the evening. A test that does not cover whole weeks compares two versions and two audiences.

The minimum rule is two complete weeks, preferably three, starting and finishing on the same day of the week. It is the only way to have every day represented equally in both groups.

That has a consequence clients find frustrating and that is better announced: a serious test occupies the site for three weeks, during which nothing else changes. A modification along the way cancels everything.

It also limits how many tests are possible in a year, and that is a good thing. A business able to run four tests a year, properly, will have learned more than one launching twenty of which none concludes.

Finally, note the start and end dates and what else was happening. A campaign launched mid-test changes the traffic mix the conversion-rate article describes, and the test no longer measures what it thought it did.

Test the offer, not the button colour

When volume is limited — which is nearly always — the only thing worth testing is what can produce a wide gap. Large gaps are detectable with few people; small ones are not detectable at all.

That eliminates almost everything the profession proposes: button colour, the position of an element, the wording of a headline. Those changes produce real but small differences, and small means undetectable at your scale.

What remains is interesting: the price, the structure of the offer, the delivery promise, the presence or absence of a price, the channel offered — conversation or form. Those are changes that move behaviour plainly.

There is a counterpart to accept: those tests are not free. Testing two prices means selling at two prices for three weeks, and that carries a real commercial cost a colour test does not.

Our position is that the cost is the price of the information. A colour test costs nothing and teaches nothing; a price test costs something and answers the one question everything else depends on.

Here, traffic moves with the calendar

One local particularity makes comparisons over time even less reliable than elsewhere, and it is why we insist on simultaneous groups rather than successive periods.

Traffic and buying behaviour move with Ramadan and Eid, with the pay week, with the start of the school year, with the regulated sales periods and with your trade’s seasons. Those movements are large, and they are larger than most of the effects you are trying to measure.

The consequence is direct: comparing this week to last week is not a test, it is a seasonal observation. We regularly see "+40%" results that describe the pay week and nothing else.

The only protection is simultaneity: both versions run at the same time, on the same traffic, split at random. The whole value of the method sits in that sentence, and it is what separates it from a before-and-after.

And for genuinely particular periods — the two weeks before Eid, for instance — the right decision is often not to test at all. What you would learn describes a situation that does not come back for a year.

What to do when you cannot test

The honest answer to insufficient volume is not to test anyway: it is to change method. And the alternative method is less prestigious and often more profitable.

The first step is a queue of obvious corrections. Missing price, number that cannot be tapped, absent delivery terms, over-long form, slow page: the conversion-rate and landing-page articles give the list. Those are not hypotheses, they are absences.

The second is direct observation. Watch five people use your site on their own phones, without helping them, and note where they hesitate. Five people are enough to find the major faults, and no statistical volume is required for it.

The third is reading the incoming messages. Every question asked about information already on the page is a page fault, and you receive a free list of them every week.

The fourth is sequential change, named as such: you modify one thing, leave it three months, and look at the trend. That is not a test and it should not be called one in a report — but it is an informed decision, and it is what most businesses can actually afford.

The subgroup trap

A test that does not conclude is often followed by a search: what if version B wins among visitors from advertising? And on phones? And in Oran? That is the moment an honest analysis becomes a fishing trip.

The problem is arithmetic. Cutting the data into ten subgroups gives ten chances of finding a gap by accident, and one gets found. It is not real, and it will be presented as the test’s discovery.

The protective rule is simple: the subgroups you will look at are decided and written down before launching, and there are few of them. Anything examined afterwards is a hypothesis for a future test, never a conclusion.

There is one legitimate exception and it is local: separating orders accepted from orders refused at delivery. That is not an exploratory subgroup, it is the correct definition of the outcome, and it has to be planned from the start.

A second exception: phone against desktop, where the site behaves differently on the two. There too it is decided beforehand, because it is a hypothesis rather than a find.

The change log

The durable deliverable of this service is not a test, it is a log. One line per entry: the date, what was changed, on which page, why, and what was observed afterwards.

Without that document a business redoes the same modifications every eighteen months, at the rhythm of supplier changes. We have seen it three times on one site: the price published, then removed, then published, each time with conviction.

The log also serves to interpret measurements. A fall in conversion in March has a different explanation if the form changed on 3 March, and that information exists nowhere else — not in the measurement tool and not in anybody’s memory.

It has to contain failures as much as successes, and above all inconclusive tests. A test that showed nothing is real information: it says the gap, if there is one, is smaller than what you can detect.

Two lines per entry is enough. What matters is that the file exists, that it lives with you, and that it survives the departure of whoever started it.

Reading a result without lying to yourself

When a test concludes, three sentences are enough to report it honestly, and they are rarely written.

The first states the question: which version, on which page, measured on which event, over exactly which period. Without that the result is not reproducible and it means nothing in six months.

The second states the gap in absolute numbers before stating it as a percentage. "Eleven accepted orders against seven" informs; "+57%" informs far less, and it is the formulation chosen precisely when the numbers are small.

The third states what remains uncertain. An honest conclusion always contains a sentence of that kind, and its absence in a report is the most reliable signal that something has been rounded in the convenient direction.

And one temptation has to be resisted: applying the winner everywhere immediately. A version that wins on an offer page does not necessarily win on the service page, because visitors do not arrive there with the same intent.

What a test really costs

The announced cost of a test is the tool’s, and that is the smallest part. The real cost spreads across four lines that never appear in a quotation.

The first is design time: deciding the question, version B, the measured event, the planned subgroups and the duration. It is half a day of serious work, and it is the part that decides whether everything else is valid.

The second is immobilisation: three weeks during which the site does not move. For a business with a queue of obvious corrections waiting, that cost is far above the test’s hoped-for gain.

The third is the commercial cost when you test what is worth testing. Two prices, two delivery promises, two offer structures: that is paid out of margin for the duration.

The fourth is technical risk: a badly installed testing tool slows the page or briefly shows version A before B. On a mobile connection that flicker is visible, and it damages both versions at once.

What we do, and what we will refuse to do

What we will refuse: launching a test on a site whose volume does not allow a conclusion. It is a frequent refusal, it is badly received, and the alternative we propose is cheaper — which does not help it be accepted.

We will refuse to stop a test because one version has pulled ahead, and to relaunch a lost test hoping for better. Both manufacture winners, and a report that always announces one describes a reading practice.

We will refuse to present a result as a percentage without the absolute numbers, and to conclude on a subgroup that was not decided before launch.

What we do: the preliminary calculation that says which side of the limit you are on; the queue of obvious corrections when the answer is no; three-full-week tests on the offer rather than the form; the result measured on the accepted order; and the change log, kept at your premises.

And what you can do this week without us: take last month’s count of accepted orders and divide it by two. Look at that number. If you would not be willing to build a decision on it, you have just learned that the next useful expense is not a test.

Frequently asked questions

How many visitors do we need to test?

There is no general answer: the volume depends on your current rate and the size of the effect you want to detect. A single figure published as the answer would be wrong for everybody. Count your conversions and divide by two instead.

Can we stop a test as soon as one version wins?

No, and it is the error that produces the most false winners. Chance creates gaps constantly; stopping at the favourable moment amounts to manufacturing proof. Fix the duration and volume before launching and do not depart from them.

How long should a test run?

Two complete weeks minimum, three preferably, starting and finishing on the same day of the week. A shorter test compares two versions and two audiences, because buying behaviour changes by day.

What should we test first?

What can produce a wide gap: the price, the structure of the offer, the delivery promise, the channel offered. A button colour produces a real difference that is too small to detect at your scale.

Can we compare this week with last week?

No. Traffic moves with Ramadan, the pay week, the school year and your trade’s seasons, and those movements are larger than the effects being sought. Both versions have to run at the same time on the same traffic.

What if we do not have enough traffic?

Correct what is missing rather than testing hypotheses: price, tappable number, delivery terms, form, speed. Then watch five people use the site on their phones — five are enough to find the major faults.

Where we come in

Last month’s orders halved put you on one side of the line or the other. Almost always the wrong side, and the question then becomes what to do instead.

  • We do the arithmetic with your figures and write the answer in one line.
  • We order your pending corrections by what each is worth.
  • We set up the dated register that stands in for the test you cannot run.

Below the volume required no test concludes: we will not sell you one, and the correction list will produce more.

Read next

Let us talk about your project

A free audit, no commitment: we look at your online presence and tell you what is holding it back.

We measure how this site is used with Google Analytics, to learn which pages actually help. You can stop that measurement at any time from the footer. Cookie policy