Backup Validation & Restore Drill Checklist
Backup-success emails do not prove recovery. This is the restore-first checklist 912 uses on managed backup engagements — weekly review, monthly restore sample, and the RPO/RTO evidence register that turns a backup policy into an auditable practice.
Built from real 912 engagements — not generic checklists.
Weekly 15-minute backup-log review
Five checks across job status, capacity, retention drift, credentials, and alert delivery. Catches silent failures before they accumulate.
Monthly restore drill — five steps, one workload
Pick a real workload, restore to an isolated target, validate at the application layer, measure real RTO and RPO, record the evidence.
RPO/RTO evidence register
One row per drill. Workload, recovery point, restore target, measured RTO, measured RPO, validator, sign-off. The audit trail vendors and insurers ask for.
Quarterly offsite restore drill
Primary-site drills prove the catalog. Offsite drills prove the DR plan. Cadence guidance for hybrid and multi-site estates.
Common failure modes flagged in advance
High-water-mark trigger thresholds, evidence-row discipline, and the offsite restore gap that audits surface most often.
What you get
A restore-first checklist for proving that backups can recover real workloads, not just complete scheduled jobs.
Files delivered
Download the Restore Drill Checklist (PDF)
Next action
Use the resource, then review the related service path: /services/backup-recovery.
The printable checklist and blank evidence register land in your inbox within 60 seconds. Use the two-page working tool across VMware, Proxmox, SQL, NAS, and cloud backup operations.
Get the Backup Restore Drill Checklist
Two-page weekly review, monthly restore drill, and blank evidence register.
Weekly cadence
The 15-minute weekly review.
Schedule this in the same calendar slot each week. Five checks; takes a quarter of an hour on a healthy estate; surfaces the warning signs before they become missed-backup weeks.
Backup job status report
Pull the last 7 days of job outcomes per protected workload — success, warning, fail, missed. A green dashboard from the backup server is not enough; review the per-workload report.
Storage target capacity check
Confirm primary and offsite targets are below the configured high-water mark. A backup that completes onto a full repository is a silent corruption risk on the next run.
Retention policy spot-check
Pick one workload at random and verify daily/weekly/monthly retention points actually exist in the catalog. Policy drift between vendor updates is the most common cause of missing recovery points.
Credential and agent health
Check that the backup service accounts have not been rotated out and that agents on protected VMs are reachable. Expired credentials show as transient warnings before they become 7 days of missed backups.
Alert-channel verification
Confirm at least one failure alert in the period reached the named on-call inbox or chat channel. Silent alerts are equivalent to no alerts.
Monthly drill
Five steps from policy to proof.
Rotate the workload each month so every protected system gets drilled at least annually. Predictable rotation removes the "we will get to it next month" failure mode.
Pick a real workload
Rotate through the protected estate: month one a file server, month two a database, month three a domain controller, month four an application server. Predictable rotation forces every workload to be drilled annually.
Restore to an isolated location
Mount or restore to a sandbox host or quarantined VLAN — never overwrite production. Document the restore target and the restore start time before the operation begins.
Validate at the application layer
Log in to the restored workload. For a file server, open recent files. For a database, run a row count on a known table. For a domain controller, query a known account. The job-success flag from the backup tool does not prove the restored state is usable.
Measure RTO and RPO against the policy
Capture elapsed time from start of restore to validated application state — that is your real RTO. Capture the timestamp of the most recent transaction inside the restored data — the delta to "now" is your real RPO. Compare both to the documented targets.
Record evidence in the register
One row per drill: workload, recovery point used, restore target, RTO measured, RPO measured, validator name, sign-off date. Without the row, the drill did not happen.
Evidence register
The row the audit asks for.
A single row per drill, captured in a spreadsheet or ticketing field. Without a written validator sign-off, the green job is not evidence — it is a status indicator.
| Column | Example entry |
|---|---|
| Workload | ERP database (anonymized) |
| Recovery point used | 2026-05-18 21:30 EAT |
| Restore target | Isolated VLAN 99, sandbox host |
| Measured RTO | 47 min (target 60 min) |
| Measured RPO | 9 hr 42 min (target 12 hr) |
| Application-layer validation | Last invoice number visible; row count matches export |
| Validator | Named IT lead |
| Sign-off date | Last business day of the month |
Common gaps
What this routine catches that vendor dashboards miss.
Mounted backup targets above 85 % usage
Most cleanup jobs cannot reclaim space fast enough during a peak retention week. Trigger storage-target remediation at 80 %, not at the failure point.
Job-success email without an evidence row
A green job and a written validator sign-off are different artefacts. If the audit only ever sees one of them, the drill did not happen.
No alternate offsite target tested
Primary-site restore proves the catalog. Offsite restore proves the DR plan. Schedule at least one offsite restore per quarter; the rest can be onsite.
Ready to talk about this in context for your business?
Schedule a strategy call