Backups that restore: a 3-2-1 plan for dedicated servers
Every team has backups. Far fewer have restores. This guide lays out a backup design for dedicated servers that holds up under the only test that matters: getting the data back, quickly, when something has gone wrong.
7 minute read
What backups are for
Mirrored drives protect against a drive failing. Backups protect against everything else: a bad deploy that truncates a table, ransomware, a deleted volume, a data centre incident, or an engineer running the right command on the wrong host. Design for those, not for disk failure.
3-2-1, applied
- Three copies: production, a local snapshot for fast rollback, and an off-site backup.
- Two media or systems: the off-site copy should not depend on the same storage system or credentials as production.
- One off-site: a different data centre at minimum, and ideally one the production credentials cannot delete.
What to back up, and how
- Databases: use the engine's own tools. For Postgres that is a base backup plus continuous WAL archiving, which gives point-in-time recovery. Never copy the data directory of a running database.
- Files and object storage: a deduplicating, encrypting tool such as Restic or Borg, or filesystem snapshots with ZFS send.
- Configuration: everything needed to rebuild the server should be in version control; the backup is the data, not the operating system.
- Secrets and keys: back up the encryption keys somewhere the backups themselves are not, or the backups are decorative.
Schedule and retention
Nightly is the floor for most systems; databases with continuous WAL archiving effectively back up every few seconds. Keep daily backups for 30 days, weekly for three months and monthly for a year if compliance asks. Retention costs storage, which on a dedicated Store server is cheap; losing the one backup from before the corruption is not.
Encryption and immutability
Encrypt at the source with keys you hold. Make the off-site copy append-only or object-locked so that the credentials on the production server can write new backups but cannot delete old ones. This single control is what turns a ransomware incident from a disaster into a restore.
The restore test
- Once a quarter, pick a backup from at least a week ago.
- Restore it to a scratch server, not to production.
- Time every step: fetch, decrypt, restore, start the database, verify row counts or checksums.
- Write down the recovery time and anything that surprised you.
- Fix the surprises before the next quarter.
Recovery objectives
Decide two numbers with the business before designing anything: how much data you can afford to lose (recovery point objective) and how long you can be down (recovery time objective). Nightly backups give a 24-hour recovery point; WAL archiving gives minutes. A restore from a second site over 10 Gbit moves a terabyte in about 20 minutes; over 1 Gbit, more than two hours.
Common failures
- The backup job silently stopped months ago and nobody alerted on its absence.
- The database backup is a file copy and will not start.
- The off-site copy is in the same account with the same credentials, and the attacker deleted both.
- The encryption key was only on the server that was lost.
- The restore works but takes three days because nobody measured it.
Packages mentioned.
- Store 300Store
308 TB of raw disk behind 15 TB of NVMe cache, run as object or block storage.
- CPU
- 32 cores / 64 threads
- Memory
- 256 GB DDR4 ECC
- Storage
- 14 × 22 TB HDD + 2 × 7.68 TB NVMe
€3,290per month - Core 48Core
48 Zen 4 cores, 256 GB and fast NVMe for your main application tier.
- CPU
- 48 cores / 96 threads
- Memory
- 256 GB DDR5 ECC
- Storage
- 2 × 3.84 TB NVMe (mirrored)
€2,540per month
Keep reading.
What it takes to run a wholesale dedicated server in production
The checklist a competent operations engineer works through between receiving a root password and trusting a dedicated server with production traffic.
Read the guideSizing a dedicated server for Postgres: cores, memory, NVMe
How to pick memory, storage and CPU for production Postgres on dedicated hardware, and why one 48-core NVMe server replaces a lot of managed database spend.
Read the guideTell us what you run. We will tell you what it costs to run it properly.
A quote within one business day, from an engineer rather than a sales script. No setup fee, three-month minimum, delivery in 48 hours.