Artificial intelligence
The pilot that will never ship: the signs, from week two
A pilot that fails rarely fails at the end. It fails in the second week, quietly, and everybody carries on for three months.
The demonstration went well. The director asked three questions, the prototype answered all three correctly, the room found it impressive, and the project was funded for three months. It will never go into service, and that was knowable by the end of the second week.
We have seen this sequence often enough to recognise its signs, and they are all early. None is technical: an artificial intelligence pilot that dies almost never dies because the model was poor. It dies because nobody answered a question that was never asked.
This article lists those signs, in the order they appear, and gives the four questions to ask in week two to settle it. These are our observations and not a study: there is no Algerian statistic on the subject, and we will not import one from elsewhere to look serious.
The demonstration worked, and that is the problem
A demonstration is prepared. The examples are chosen, the order is rehearsed, and the person giving it knows the questions that will come because they are the ones they prepared for. None of that is dishonest; it is what a demonstration is.
The problem is what it produces in the room: the conviction that the hard part is done. It has just finished the visible part, which is also the easiest, and the room has just committed a budget on that basis.
The useful reversal is to treat a successful demonstration as teaching nothing. It establishes that the system can work on chosen cases, which was likely, and establishes nothing about the unchosen ones, which are the bulk of the real work.
What does teach something is the opposite: give the prototype fifty cases taken at random from what actually happened last month, discarding none, and watch what it does with them in front of the people who handled them. That session is uncomfortable and it is the only one that informs.
What a pilot should prove, and what it proves instead
A pilot has one function: to reduce a named uncertainty. Before launching one, somebody has to be able to say the sentence “we do not know whether…” and finish it. Without that sentence the pilot reduces nothing, it occupies.
What it proves instead, almost always, is that the technology exists. That is information nobody needed: it is available free, in ten minutes, on any consumer tool. Three months of a team to obtain it is the commonest waste in this field.
The uncertainties that deserve a pilot are of a different kind and they are always local. Does the system understand our customers when they write the way they write. Is our data clean enough for it to rely on. Will our teams look at it. None of the three is settled by a demonstration.
The test to apply before launching fits in one question: if the pilot succeeds, what will we know that we do not know today? A vague answer here is a project that will not end, and knowing it costs five minutes instead of three months.
Week two: nobody has provided real data yet
This is the first sign and the most reliable of the early ones. By the end of the second week the prototype is running on invented data, on last year’s export, or on thirty lines copied by hand. Nobody has delivered what the technical team asked for on day one.
It is almost never obstruction. It is that extracting real data requires a decision nobody has taken: who is entitled to extract it, in what form, with which fields removed, and who signs. Those four questions are exactly a production project’s questions, and the pilot has just hit them two weeks in.
The sign does not say the project is bad. It says the organisation is not ready to put it into service, because going live will require the same decision at a larger scale. A pilot that cannot obtain thirty real lines will not obtain a daily flow.
The useful response is to treat that as the subject rather than as a delay. Stopping development for a week and settling the extraction question — what the register describes, field by field is the same information — beats carrying on with wrong data for two months.
The example set never changes
Second sign, and it is visible to the eye. The same six or seven examples have circulated since day one: they are in the demonstration, in the screenshots, in the technical team’s Thursday message. Nobody has added to them in two weeks.
An example set that does not grow means nobody outside the technical team is trying the system. And that is where the cases that matter live: business people do not bring typical examples, they bring the ones that annoy them, and those are what decide whether the system is useful.
There is a harder variant to spot: the set grows but always from the same person. That produces a system excellent at the way that person asks questions and mediocre at everybody else’s, which is only discovered at go-live.
The remedy is trivial and it is nearly always refused for bad reasons: put the prototype in the hands of three business people in week two, ugly, incomplete and embarrassing as it is. A prototype nobody dares show is a prototype that will learn nothing.
The success criterion was never written down
Third sign. Ask three people on the project what would make you say the pilot had succeeded. If you get three different answers, or three that start with “if it works well”, there is no criterion, and a pilot without one cannot end.
A useful criterion has a recognisable shape: a number, a threshold and a population. “On a hundred real delivery enquiries, the system answers seventy alone and correctly, and hands the other thirty over without error” is a criterion. “The system is reliable” is not.
The absence of a criterion has a consequence that explains most pilots that drag: with no threshold, no result allows anybody to say no. Each disappointment becomes an adjustment, each adjustment takes two weeks, and the project becomes continuous development nobody ever decided on.
Writing the criterion after the start is possible and honest, on one condition: write it before looking at the week’s results. A threshold set while already knowing the score is not a threshold, it is a justification.
The business team has not opened the screen
Fourth sign, and it is measurable without asking anybody: look at who logged in. If the only sessions on the prototype come from the technical team and the person who launched the project, the pilot is running in a closed loop.
The reason is rarely lack of interest. It is more often that opening the screen means leaving the tool those people work in all day, to try something that is not yet their job. A prototype living in a separate tab is a prototype opened when somebody remembers it, which is to say not.
It is also the best predictor of what happens after go-live, and our observation on this point is consistent: a team that did not open the screen during the pilot will not open it more once it is official. Making something official creates no habit, it only makes the absence of one visible.
The correction is to bring the prototype to where the work happens rather than the reverse — into the messaging they use, into the tool they already have open, onto the phone. That is more work in week two and it is what decides everything else.
The hard cases are deferred to phase two
Fifth sign, and it is verbal: listen for whether “we will look at that in phase two” recurs. Note what it covers. In projects that finish, it covers refinements; in those that do not, it covers the cases that make up most of the volume.
Deferring to phase two is a comfort mechanism and it works very well: it lets the pilot keep working nicely. It moves the difficulty to a moment when the budget is committed, the team is tired, and the gap between what is handled and what was supposed to be handled becomes impossible to absorb.
There is a particular case for systems that answer: the requests the prototype sets aside. They are normal and they have to be counted from week two, because the pile of what is not handled needs somebody who opens it every morning — and that is a person, not a line of code.
The rule we apply is to take the hardest case first, and it is unpopular for an obvious reason: it makes the demonstration less attractive. In exchange it gives the only information worth having, which is whether the project is possible at all.
The running budget exists nowhere
Sixth sign. The pilot has a budget, the team has a budget, and nobody has written what the system will cost per month once in service. That is not an accounting oversight: it is the symptom of a project not yet thought of as a service.
The lines are known and list in ten minutes: model usage, hosting, maintenance, and the human work the system does not do. The threshold between hosting and calling depends on figures that do not exist yet in week two, and that is fine — what matters is that the line exists, not that it is right.
The absence of that line predicts something specific: at go-live somebody will discover a monthly cost that was in no table, and the decision will be reopened in a context where reopening it is expensive. We have seen it stop technically successful projects.
The good practice is a rough figure written in week two, to an order of magnitude, with its date. It will mostly serve to be corrected, and its real function is to force the question “and then what” while it is still cheap to ask.
The most reliable sign: the question that never comes
Here is the one we look at before all the others, because it summarises them. In three weeks of meetings, has anybody asked who will use this system on an ordinary Tuesday, and who will deal with it when it gets something wrong?
In projects that finish, that question arrives early and it arrives from the business rather than from technology. It has a recognisable shape: it is about a person and a moment, not about a capability. “Who will answer when it hands over?” is that question; “can it write summaries?” is not.
When it is never asked, it is almost always because the project is carried by interest in the technology rather than by a need. That is not a fault, and it sometimes produces very good learning, provided it is named — an exploration project has value, and it has no go-live.
The corollary is harder: if nobody asks the question, it will not be asked after go-live either. The system will be delivered to an organisation that was not waiting for it, and it will join the list of tools the business owns and does not open.
What happens when it is allowed to run
The pilot does not stop, and that is what makes it expensive. It becomes a permanent demonstration, brought out when a visitor passes, improved a little before every committee, occupying part of a team for months without ever changing state.
The second cost is less visible and more lasting: it consumes the subject’s credibility inside the business. After two pilots that produced nothing, the third proposal, however good, is received by people who have seen this film. We regularly meet businesses whose real obstacle is a project from two years ago.
The third is human. The people who worked on a project that does not finish generally know, well before their management, that it will not, and they spend weeks doing something they have stopped expecting a result from. It is an expensive way to occupy capable people.
Stopping early is therefore the decision that preserves the most, and it is also the hardest to take, because a pilot working nicely on its own examples does not look like a failure. That is exactly what makes the early signs useful: they are visible before the demonstration becomes the product.
The week-two check: four questions
Block out thirty minutes at the end of the second week and ask these four questions, in this order, of the people on the project — without preparing the meeting and without announcing the questions in advance.
One: what data is the system running on right now, and where exactly did it come from? Two: how many people outside the technical team opened it this week, and how many new examples arrived? Three: what number, on what population, will make us say it succeeded? Four: who will use it on an ordinary Tuesday, and who will handle what it sets aside?
The reading rule is simple and worth sticking to: two vague answers out of four, and the pilot is already in the process of not finishing. That is not a prediction, it is a description of its present state — none of the four is about the future.
The check needs no tool and no supplier, takes half an hour, and its commonest result is to save two months. It is the one thing in this article we ask you to do without us, and it is the one with the most value.
What we do, and what we refuse
We scope a pilot around a written uncertainty and a numbered threshold, with real data from the first week or no pilot at all. We put the prototype into the tool your teams already open rather than into one more screen, and we take the hardest case first, which makes the demonstration less enjoyable and the result usable.
We ask the four questions above at the halfway point, including when the answer costs us the rest of the engagement. It is the part of the work that has made us stop projects we could have extended, and it is why our recommendations on the projects we continue are worth anything.
We do not deliver a demonstration built to win over a committee. We are asked to, the request is understandable, and it produces exactly the mechanism described in this article: a room convinced by chosen examples, a budget committed on that basis, and a team discovering the real cases three months later.
And we promise no success rate before having seen your data. It depends on how clean it is and on the variety of your cases, two things we discover at the same time you do. A figure announced before looking is precisely the kind of figure that starts the pilots this article is about.
Frequently asked questions
How long should a pilot run?
Long enough to answer the uncertainty written at the start, and no longer, which in practice means a few weeks rather than a few months. A duration fixed in advance with no criterion produces a project that stops when the calendar says so, which is not the same as concluding; a criterion with no duration produces the opposite. You need both.
Do we really need real data from the start?
You need real data, cleaned if necessary, and not invented data. A fabricated set contains the cases whoever fabricated it thought of, which is to say the easy ones, and a pilot succeeding on it has measured nothing. Cleaning or anonymising is legitimate work; inventing is not.
Our pilot is working well. Is that still a bad sign?
It depends entirely on what it is working well on. If it works on cases brought by business people, chosen by them and not by the technical team, that is the best sign there is. If it works on day one’s example set, it only tells you that set is handled well.
Can a stopped pilot be restarted?
Yes, and it is often easier than a first attempt, provided you pick up what was missing rather than what was built. In nearly every case we have seen, what was missing was not technical: the data, the criterion, or the person who was going to use it. The code is quick to rewrite.
Who should carry a pilot in a small business?
Somebody from the trade concerned, with time genuinely freed to open it several times a week. A pilot carried by senior management advances through meetings and one carried by IT advances without users; neither produces the only thing that matters, which is use.
Are these signs taken from a study?
No, and we prefer to say so: they are our observations from projects we have run and from ones clients have described to us. There is no Algerian statistic on failure rates for this kind of project, and a foreign study quoted here as a figure would describe another market with the authority of a number.
Where we come in
Two vague answers to the four week-two questions, and the question is no longer whether the pilot will finish — it is how many months it will occupy before that is admitted.
- We set down in writing the question the trial has to answer, and the number that will answer it.
- The trial lives where your teams already work — their messaging, their usual screen — and nowhere else.
- We begin with whatever resists, in front of the people who deal with it today.
Nothing gets built to impress a room, and a percentage announced without opening your files would not be one.
Read next
Custom software: the specification you write is not the one you need
Nobody knows what they want before using it. A complete specification written beforehand is a guess set down in the present tense.After version one: what evolving custom software really costs
Custom software is not delivered, it is put into service. Everything that matters afterwards turns on how you ask for a change.The clause your supplier contract does not contain
Your contract describes a service: availability, support, price. It almost never says what the supplier is allowed to do with your data.
Let us talk about your project
A free audit, no commitment: we look at your online presence and tell you what is holding it back.