Skip to main content

Backup and Disaster Recovery

Use when defining RPO/RTO or backup/restore systems.

Overview​

Design and verify backup systems from recovery objectives backward. A backup is not evidence of recoverability; only a timed, integrity-checked restore of the complete service is. Keep policy decisions fail-closed until provider, retention, storage, encryption, key custody, automation, and restore evidence are all approved.

When to Use​

  • Defining or changing RPO and RTO
  • Selecting managed backups, PITR, replication, or external object storage
  • Writing restore, rollback, roll-forward, or incident runbooks
  • Automating database or object backups
  • Testing clean migration chains and upgrade-path recovery
  • Comparing low-cost disaster-recovery options

Core Distinctions​

RPO is acceptable data loss​

An RPO of 24 hours requires a recoverable point no more than 24 hours old. Measure backup completion and recoverability, not merely when a scheduled job started. Alert on stale age, missing objects, failed encryption, integrity mismatch, and failed verification.

RTO is complete service restoration​

RTO starts at the approved incident-declaration boundary and ends only when the service is safe for users. Include:

  1. Authorization and writer freeze
  2. Backup discovery, download, decryption, and integrity verification
  3. Infrastructure provisioning or target selection
  4. Schema, data, and security-object restoration
  5. Migrations and compatibility handling
  6. Secrets, endpoints, DNS, workers/functions, and object-store reconfiguration
  7. Referential, checksum, authorization-denial, readiness, and smoke tests
  8. RPO-gap reconciliation and controlled reopening of writes

Provider restore duration alone is not an RTO measurement.

Objective, evidence, and guarantee differ​

  • Objective: the target chosen by the owner.
  • Evidence: timed drills showing the current system can achieve it.
  • Guarantee/SLA: a contractual provider commitment.

Never present an objective or one successful drill as a provider guarantee. If restore time varies with data size, define a tested size ceiling and requalify after material growth.

Procedure​

1. Inventory recovery scope​

List every stateful class and its recovery mechanism:

  • Database schemas, data, migration metadata, roles, grants, RLS/policies, extensions, triggers, and security-definer functions
  • Identity/auth data and custom-role credentials
  • Object-storage objects and bucket configuration; database backups may contain only object metadata
  • Application secrets, key references, worker/function configuration, DNS, and external connector state
  • Audit, approval, operator, and deletion-propagation evidence

Name exclusions explicitly. A database-only plan is incomplete when files or external state exist.

2. Make the policy contract fail closed​

Represent each production choice explicitly:

  • RPO and RTO
  • Primary backup provider/path
  • Independent-copy destination
  • Retention and deletion propagation
  • Encryption boundary and recovery-key custody
  • Restore authorization
  • Drill frequency and pass thresholds

A partially chosen policy stays inactive. It is valid to approve RPO/RTO while leaving the provider policy unselected.

3. Compare options using current primary sources​

For every candidate, verify current official documentation and pricing. Compare:

DimensionQuestions
RecoveryDoes it meet RPO granularity? Can it restore in place and to an isolated target?
ScopeDatabase only, or auth/files/config too? What is omitted?
TimeIs restore duration bounded, size-dependent, or covered by an SLA?
IndependenceIs there a copy outside the primary database account/provider boundary?
SecurityEncryption, immutable retention, least privilege, key separation, auditability?
CostBaseline platform cost vs incremental backup cost; storage, operations, egress, compute, standby, and drill costs?
OperationsScheduling reliability, stale-backup alerts, credential rotation, and restore complexity?

Use the arithmetic tool for cost scenarios. Show assumptions and distinguish total platform cost from incremental backup cost. Re-check vendor facts before every decision; pricing and retention change.

4. Prefer layered recovery​

For small systems, a common low-cost shape is:

  1. Managed daily/continuous provider backup as the primary path
  2. Client-side-encrypted logical export to independent private object storage
  3. Immutable object names plus retention lock/lifecycle
  4. Automated freshness and integrity checks
  5. Regular isolated restore drills

The independent export is defense in depth, not permission to weaken the managed path. A third-party scheduler must not be the sole RPO control unless its timing and failure behavior are qualified.

5. Design encryption and authority boundaries​

  • Encrypt before upload when backups contain sensitive data.
  • Keep the decryption key outside the repository, scheduler, storage provider, and backup payload.
  • Ensure authorized responders can retrieve the key during an incident; untested offline custody creates an unusable backup.
  • Use a dedicated least-privilege backup principal.
  • Separate automated write credentials from restore/read and retention-policy administration.
  • Never print connection strings, tokens, keys, or dump content.
  • Delete plaintext and staging artifacts on success and failure.
  • Store a manifest with migration head, ciphertext size, and cryptographic checksum or signature.

6. Define retention and immutability together​

Retention is not just lifecycle deletion. Protect recovery points against overwrite and premature deletion with object lock, bucket lock, or equivalent. Align retention with privacy deletion propagation: document when deleted user data ages out and how emergency restores avoid silently resurrecting it.

7. Build a timed restore gate​

Reserve part of RTO for incident authorization and application validation. For a one-hour RTO, a reasonable starting budget is 45 minutes for technical restore and 15 minutes for checks and reopening, but the owner must approve the split.

A drill passes only if it:

  • Uses production-sized or explicitly bounded representative data
  • Starts from an isolated clean target
  • Replays the clean migration chain and exercises the upgrade path where applicable
  • Verifies row counts, referential integrity, canonical checksums, migration versions, and security catalog
  • Exercises runtime behavior under non-owner roles
  • Restores or explicitly reconfigures auth, files, secrets, endpoints, and workers/functions
  • Cleans temporary dumps and resources on success and failure
  • Records stable machine-readable evidence without secrets
  • Completes end to end within RTO

Requalify after material data growth, schema/security changes, provider changes, failed backups/drills, and at the approved recurring cadence.

8. Define escalation before the drill fails​

Pre-approve the decision ladder:

  1. Optimize dump/restore and remove avoidable manual steps.
  2. Add a provisioned recovery target or warm standby.
  3. Move to PITR/continuous archiving or replication.
  4. Buy a provider SLA/support tier if contractual assurance is required.

Do not claim the RTO while the timed gate is red.

Automation and Alerts​

Alert on:

  • Job failure or missed schedule
  • Backup age beyond policy
  • Empty or unexpectedly small/large output
  • Encryption or manifest failure
  • Upload/read-back failure
  • Integrity mismatch
  • Retention-lock/lifecycle drift
  • Restore-drill failure or duration regression

Test the alert path; a silent scheduled job is not an operational control.

Pitfalls​

  • Choosing numbers but no system: RPO/RTO values alone do not activate a recovery policy.
  • Calling daily scheduling proof of 24-hour RPO: completion, freshness, integrity, and recoverability must be monitored.
  • Equating provider restore with service recovery: configuration, secrets, files, auth, validation, and reopening consume RTO.
  • Backing up only PostgreSQL: object files and external configuration may be omitted.
  • Provider-managed encryption only: it may not satisfy key-separation requirements for sensitive exports.
  • One credential can write, read, delete, and change retention: compromise destroys production and recovery evidence.
  • Testing against tiny fixtures only: a size-dependent restore needs production-sized timing evidence.
  • Publishing precise cost without assumptions: calculate retained GB-month, operations, runner time, compute, and standby separately.
  • Selecting a provider before approval/drill: document it as a candidate and keep the production contract fail closed.

Verification Checklist​

  • RPO and RTO and their measurement boundaries are explicit.
  • Every stateful data/configuration class has a recovery path or explicit exclusion.
  • Current official docs and pricing support provider and cost claims.
  • Baseline and incremental costs are separated with assumptions.
  • Backup credentials and recovery keys follow least privilege and separation of duties.
  • Independent copies are encrypted, immutable for the retention window, and lifecycle-managed.
  • Freshness, integrity, policy drift, and restore failures alert.
  • A production-sized isolated restore completed within RTO and produced evidence.
  • Rollback versus roll-forward rules and incident authorization are documented.
  • Privacy deletion propagation through backups is documented.
  • The production policy remains inactive until every required decision and gate passes.

References​

  • For a dated, source-backed managed PostgreSQL plus object-storage cost pattern, see references/managed-postgres-low-cost-pattern.md. Re-verify every vendor fact before reuse.

Supporting files: this skill's supporting files are held in the docsite at docs/15-skills/_support/engineering/backup-and-disaster-recovery/ — fetch them fresh from jknash/docsite main alongside this page. Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/engineering/backup-and-disaster-recovery/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

version 1.0.0.

Published by Muse · 2026-10-04.