Skip to content

Backups that restore: a 3-2-1 plan for dedicated servers

Every team has backups. Far fewer have restores. This guide lays out a backup design for dedicated servers that holds up under the only test that matters: getting the data back, quickly, when something has gone wrong.

7 minute read

What backups are for

Mirrored drives protect against a drive failing. Backups protect against everything else: a bad deploy that truncates a table, ransomware, a deleted volume, a data centre incident, or an engineer running the right command on the wrong host. Design for those, not for disk failure.

3-2-1, applied

  • Three copies: production, a local snapshot for fast rollback, and an off-site backup.
  • Two media or systems: the off-site copy should not depend on the same storage system or credentials as production.
  • One off-site: a different data centre at minimum, and ideally one the production credentials cannot delete.

What to back up, and how

  • Databases: use the engine's own tools. For Postgres that is a base backup plus continuous WAL archiving, which gives point-in-time recovery. Never copy the data directory of a running database.
  • Files and object storage: a deduplicating, encrypting tool such as Restic or Borg, or filesystem snapshots with ZFS send.
  • Configuration: everything needed to rebuild the server should be in version control; the backup is the data, not the operating system.
  • Secrets and keys: back up the encryption keys somewhere the backups themselves are not, or the backups are decorative.

Schedule and retention

Nightly is the floor for most systems; databases with continuous WAL archiving effectively back up every few seconds. Keep daily backups for 30 days, weekly for three months and monthly for a year if compliance asks. Retention costs storage, which on a dedicated Store server is cheap; losing the one backup from before the corruption is not.

Encryption and immutability

Encrypt at the source with keys you hold. Make the off-site copy append-only or object-locked so that the credentials on the production server can write new backups but cannot delete old ones. This single control is what turns a ransomware incident from a disaster into a restore.

The restore test

  1. Once a quarter, pick a backup from at least a week ago.
  2. Restore it to a scratch server, not to production.
  3. Time every step: fetch, decrypt, restore, start the database, verify row counts or checksums.
  4. Write down the recovery time and anything that surprised you.
  5. Fix the surprises before the next quarter.

Recovery objectives

Decide two numbers with the business before designing anything: how much data you can afford to lose (recovery point objective) and how long you can be down (recovery time objective). Nightly backups give a 24-hour recovery point; WAL archiving gives minutes. A restore from a second site over 10 Gbit moves a terabyte in about 20 minutes; over 1 Gbit, more than two hours.

Common failures

  • The backup job silently stopped months ago and nobody alerted on its absence.
  • The database backup is a file copy and will not start.
  • The off-site copy is in the same account with the same credentials, and the attacker deleted both.
  • The encryption key was only on the server that was lost.
  • The restore works but takes three days because nobody measured it.

Tell us what you run. We will tell you what it costs to run it properly.

A quote within one business day, from an engineer rather than a sales script. No setup fee, three-month minimum, delivery in 48 hours.