IT infrastructure
The restore test: the one day you know the backup exists
A backup never restored is a hypothesis. The test is scheduled, timed and recorded — twice a year, onto a machine that is not the original.
The article accompanying this one sets out the three-copy rule, says what installations forget to back up, and devotes a section to the claim that matters most: restoration is the only proof. Then it moves on, because it has eleven other subjects.
This page is that section carried out. A restore test is a dated operation, with a responsible person, a destination machine, a stopwatch and a one-page report.
The distinction to hold from here on: a backup that runs and a backup that restores are two different claims, and the first is verified every day by a green light while the second is verified by nobody.
We give no reference duration anywhere in this article, and the reason is stated in the section it belongs to: the only useful figure is the one your own test produces, and quoting one in advance would hand you a reason not to measure it.
What is tested, and what people think is tested
Most businesses that say they test their backups restore a file. That is better than nothing and proves far less than the conclusion drawn from it.
Restoring a file proves the medium is readable and the software works. It does not prove the set is coherent, that the database starts, that permissions follow, or that somebody could do it on a Tuesday morning at eight during an outage.
The question the test has to answer is wider, and it is phrased like this: if this machine disappeared this morning, how long before the business is working again, and what would be missing when it is?
That means restoring a system rather than a file — at minimum the machine or server carrying the business, with its data, in a state somebody can work in.
The phrasing that helps, and the one we use in reports: a green light says a copy was written. It says nothing about whether it can be read back, and those are two distinct claims that everyday use conflates.
Twice a year, and which two
Two full tests a year are enough for most businesses of this size, and regularity matters more than frequency.
The two dates go into a shared calendar at the start of the year, at quiet moments in your business rather than ours. A test scheduled for "when we have a moment" never happens — it is the same mechanism as any undated task.
To those are added light checks that are not tests and must not be counted as such: opening a restored file at random once a month, and looking at whether the logs report a failure. Five minutes, and it catches outright breakage.
A third test is required outside the calendar whenever something structural changes: a server replaced, a migration, a change of backup software, a new site. Those moments are precisely when a backup silently stops covering what it covered.
A note on the Algerian calendar with practical weight: avoid periods when the team is thin. A test goes better when the people needed to say "yes, that works" are present.
Who runs it, and why not the person who installed it
The test has to be run by somebody who did not configure the backup, and it is the most neglected rule here because it is the least convenient.
The reason is not distrust. Whoever built the system knows the paths, the passwords and the workarounds; they therefore restore by leaning on things that exist only in their head, and the test succeeds for a reason that will not be available on the day of the failure.
On the day of the failure the person at the screen is often somebody else — a colleague, a different provider, you. The test should resemble that day.
The practical form is simple: whoever installed it stays reachable but does not touch the keyboard, and every time they have to be called, it is noted. The number of calls is a result of the test just as much as the duration.
In a very small organisation where one person knows how to do everything, the rule becomes: that person writes the procedure beforehand and somebody else executes it by following it. If it is not sufficient, the procedure is the missing deliverable.
Restore where: never onto the original machine
The destination of the test is the most important technical decision on this page, and getting it wrong turns a check into an incident.
Restoring onto the original machine overwrites the current state. If the backup is incomplete — which is exactly what the test is trying to discover — you have just replaced live data with a partial copy, and there is no original left to compare against.
This section is not the companion article’s three-copy rule, and the difference is worth stating. That rule concerns where copies live; this one concerns where a restore is performed. You can respect the first perfectly and destroy your data by ignoring the second.
Acceptable destinations are a spare machine, a virtual machine created for the occasion, or a new disk on loan hardware. None is expensive, and a disposable virtual machine is the most convenient: it is deleted at the end of the test.
One precaution costing a minute that avoids an incident: disconnect the test machine from the production network before starting it. A restored server booting in the belief that it is the original may claim an address, a role or a session that is already occupied.
The stopwatch is the result
The test produces a number, and that number is the reason for doing it. Without it, all that remains is an impression.
Start the clock at the moment the decision to restore is taken, not when the copy begins to be written. The minutes spent finding the right medium, the right password and the right procedure are part of the outage and are often half the total.
Stop it when somebody whose job it is has opened the system and said the business can resume. Not when the software says "complete": between the two there are checks, connections to remake and sometimes an unpleasant discovery.
That number is then compared against the only thing that matters, which is the outage the business declared it could bear. The companion article calls it the first of the two numbers everything else derives from; here it stops being an intention.
We publish no reference duration, and that is deliberate. The figure depends on your volume, your hardware and your link; more to the point, a reader given a typical duration has been handed a reason not to measure their own, and the measurement is all this article asks for.
What is restored, and in what order
The list is prepared before the test, and the order is the one in which the business restarts rather than the one in which the machines are racked.
Start with what blocks everything else: the directory or controller handing out accounts, then the file server or database carrying the business, then the workstations. Restoring a workstation before the directory produces a machine nobody can log into.
Then go down a level to what does not live on servers: mail, the accounting documents, the files stored on the machine of the person who keeps them. The companion article gives a whole section to what installations forget; the test is where that forgetting becomes visible.
Do not leave out the items that are not data: the firewall configuration, the switch configuration, the licences, the list of accounts. A business whose data comes back perfectly and where nobody can reconfigure the network has not restarted.
Finally, decide in advance how far the test goes. A test claiming to restore everything takes two days and is never repeated; a test restoring the critical chain takes half a day and happens twice a year.
The missing password, and the other dependencies
The test rarely fails on the data. It fails on the things around it, and they are always the same ones.
The administrative password for the backup system itself, known to one person. The encryption key, kept on the system being restored. The hosting provider’s account, opened in the name of an employee who has left.
Then come the external dependencies: the domain name and its records, the licence requiring online activation, the business application whose publisher has to supply a key, the internet link itself. A perfect restore of a management server is worth nothing if nobody can reactivate the software before Monday.
The remedy is a one-page document, kept outside the backed-up system, listing those items and who holds them. It does not need to be sophisticated; it needs to exist somewhere other than the machine that will have burned.
The test is the only moment that missing document is discovered without consequence. That is the main reason to run one, before even the stopwatch: the list of dependencies is only written correctly once somebody has tried to do without them.
The partial restore, which is what actually gets used
Total loss is the scenario everybody talks about. The frequent case is much smaller, and it has to be tested separately because it fails differently.
Somebody deletes a folder on Thursday and notices on Monday. An invoice is overwritten by a wrong version. A mailbox loses three months. Those requests arrive several times a year in a twenty-person business, and total loss almost never arrives at all.
What that case demands and the large scenario does not is depth of history: being able to go back to last Thursday and not merely to last night. A backup that overwrites yesterday’s copy protects against fire and not against human error, which is the more frequent risk.
So test both, and test the partial restore more often: ask once a quarter for a specific file in a specific state — "this document as it was six weeks ago" — and see whether it comes back and how long it takes.
That request has a second useful effect: it checks that whoever performs it can do so without preparation, which is exactly the situation they will be in.
What the test almost always reveals
Four findings recur, and knowing them in advance stops anybody treating one as an isolated incident.
The first is a volume larger than expected. The backup covers what it was pointed at perfectly well, and the business has since created a share, a folder or an application nobody thought to add.
The second is a duration two to three times what the team stated. That is not negligence: the estimate covered copying the data, and the total includes searching, configuring, checking and hesitating.
The third is a missing dependency, almost always one of the three in the previous section. The fourth is more embarrassing and just as ordinary: a machine discovered not to have been backed up at all, usually because it was added after the installation.
None of those four is a failure of the test. They are its results, and a test producing none of them was probably run by the person who configured everything.
The report fits on one page
With no written trace the test happened and the business keeps nothing from it. The page is written the same day, while the details are fresh.
Six lines are enough: the date, who ran it, what was restored, the measured duration, what was missing, and what is decided with a deadline. Nothing else is read six months later.
The decisions line is the only one that produces an effect. "Add the accounting share to the backup, before the 30th" is worth more than three paragraphs of analysis, and it is checked at the next test in ten seconds.
Keep the reports together, and outside the backed-up system. Two years of these pages show something no single one shows: whether the duration is falling, whether the same gaps recur, and whether decisions get applied.
It is also the document an insurer, an important client or an auditor will one day ask for. Better to take it out of a folder than to write it after the fact, and a dated report is evidence where an assertion is not.
When the test fails: what to do that week
A failed test is good news arriving at the right moment, and the only possible mistake is treating it as bad.
The first thing to do is not to reconfigure in haste. A backup modified within the hour after a failure, under the pressure of worry, is a backup nobody knows the coverage of any more.
Write down first what failed exactly, at what stage, and what would have been lost had it been real. That sentence decides the urgency, and it is very different depending on whether what was missing was a shared folder or the database carrying invoicing.
Then correct one thing at a time, and rerun the test on the corrected scope — not the full test, the part that failed. A full test rerun immediately costs a day and teaches nothing further.
And fix the date of the next full test before closing the subject. That is where the discipline is lost: the failure is corrected, everybody is relieved, and the next test is never rescheduled.
What we do, and what we refuse
What we do is bounded. We prepare the list and the restore order, we supply the destination machine, we hold the stopwatch, we write the one-page report, and we come back six months later to check the decisions were applied.
We refuse to run the test alone on infrastructure we installed ourselves. It is the conflict described in section three: we know the paths and the passwords, so we would pass a test the real failure would fail. Somebody of yours holds the keyboard, or we ask for a second provider.
We also refuse to sign a report carrying the words "restore successful" with no measured duration beside it. Without the figure the phrase is an impression, and it serves precisely to avoid looking at the one result that commits anybody to anything.
And there is one thing to do this week without us, which decides the rest: ask whoever manages your backups to restore a specific document in the state it was six weeks ago, without preparation, and note how long it takes. If the answer is "let me look into it", you have already learned what this article had to teach you.
Frequently asked questions
How often should we test?
Two full tests a year are enough for most businesses of this size, plus a light monthly check — open a restored file at random, look at the logs. An extra test is required outside the calendar at every structural change: a server replaced, a migration, new backup software. Those are precisely the moments a backup silently stops covering what it covered.
Can we restore onto the original machine?
No, and it is the mistake that turns a check into an incident. If the backup is incomplete — which the test exists to discover — you have just overwritten live data with a partial copy, with no original to compare. Use a spare machine, a new disk, or a disposable virtual machine, and disconnect it from the production network before starting it.
Why should the person who installed it not run the test?
Not through distrust: they know the paths, the passwords and the workarounds, so they restore by leaning on things that exist only in their head. On the day of the failure the person at the screen is often somebody else. They stay reachable but do not touch the keyboard, and the number of times they have to be called is a result of the test just as much as the duration.
How long should a restore take?
We publish no reference duration, and that is deliberate. It depends on your volume, your hardware and your link — but more to the point, a reader given a typical duration has been handed a reason not to measure their own. The only useful figure is the one your test produces, compared against the outage you declared you could bear.
What exactly is being tested?
A system, not a file. Restoring a file proves the medium is readable; it does not prove the database starts, that permissions follow, or that somebody could do it on a Tuesday at eight. Test the partial restore separately too — a specific document in its state six weeks ago — because that is the case that genuinely happens several times a year.
The test failed. What now?
Do not reconfigure within the hour: a backup modified under the pressure of worry is one nobody knows the coverage of any more. Write down what failed, at what stage, and what would really have been lost. Correct one thing at a time, rerun the test on the corrected part only, and fix the date of the next full test before closing the subject.
Where we come in
A restore request sent without warning is the only test worth running, and it happens this week. What gets prepared afterwards is the day it is no longer a test.
- We establish the order in which things must come back, before they are needed.
- We bring a destination machine so nothing touches production.
- We write the measured duration beside the word “successful”, which alone says nothing.
Us inspecting our own installation has no value here: ask a third party for that test, and keep us for the preparation.
Read next
Backups and continuity: what actually gets a business running again
A backup that has never been restored is not a backup. It is an assumption.Cloud hosting: what costs is not the storage
The line everybody compares is the cheapest on the invoice. The other three get paid on the day you need them most.The account, not the server: what you actually lose in the cloud
Everything you keep there passes through one login, one mailbox and one card. What breaks that chain, and in what order.
Let us talk about your project
A free audit, no commitment: we look at your online presence and tell you what is holding it back.