Backup and Disaster Recovery
Use when defining RPO/RTO or backup/restore systems.
Overview
Design and verify backup systems from recovery objectives backward. A backup is not evidence of recoverability; only a timed, integrity-checked restore of the complete service is. Keep policy decisions fail-closed until provider, retention, storage, encryption, key custody, automation, and restore evidence are all approved.
When to Use
- Defining or changing RPO and RTO
- Selecting managed backups, PITR, replication, or external object storage
- Writing restore, rollback, roll-forward, or incident runbooks
- Automating database or object backups
- Testing clean migration chains and upgrade-path recovery
- Comparing low-cost disaster-recovery options
Core Distinctions
RPO is acceptable data loss
An RPO of 24 hours requires a recoverable point no more than 24 hours old. Measure backup completion and recoverability, not merely when a scheduled job started. Alert on stale age, missing objects, failed encryption, integrity mismatch, and failed verification.
RTO is complete service restoration
RTO starts at the approved incident-declaration boundary and ends only when the service is safe for users. Include:
- Authorization and writer freeze
- Backup discovery, download, decryption, and integrity verification
- Infrastructure provisioning or target selection
- Schema, data, and security-object restoration
- Migrations and compatibility handling
- Secrets, endpoints, DNS, workers/functions, and object-store reconfiguration
- Referential, checksum, authorization-denial, readiness, and smoke tests
- RPO-gap reconciliation and controlled reopening of writes
Provider restore duration alone is not an RTO measurement.
Objective, evidence, and guarantee differ
- Objective: the target chosen by the owner.
- Evidence: timed drills showing the current system can achieve it.
- Guarantee/SLA: a contractual provider commitment.
Never present an objective or one successful drill as a provider guarantee. If restore time varies with data size, define a tested size ceiling and requalify after material growth.
Procedure
1. Inventory recovery scope
List every stateful class and its recovery mechanism:
- Database schemas, data, migration metadata, roles, grants, RLS/policies, extensions, triggers, and security-definer functions
- Identity/auth data and custom-role credentials
- Object-storage objects and bucket configuration; database backups may contain only object metadata
- Application secrets, key references, worker/function configuration, DNS, and external connector state
- Audit, approval, operator, and deletion-propagation evidence
Name exclusions explicitly. A database-only plan is incomplete when files or external state exist.
2. Make the policy contract fail closed
Represent each production choice explicitly:
- RPO and RTO
- Primary backup provider/path
- Independent-copy destination
- Retention and deletion propagation
- Encryption boundary and recovery-key custody
- Restore authorization
- Drill frequency and pass thresholds
A partially chosen policy stays inactive. It is valid to approve RPO/RTO while leaving the provider policy unselected.
3. Compare options using current primary sources
For every candidate, verify current official documentation and pricing. Compare:
| Dimension | Questions |
|---|---|
| Recovery | Does it meet RPO granularity? Can it restore in place and to an isolated target? |
| Scope | Database only, or auth/files/config too? What is omitted? |
| Time | Is restore duration bounded, size-dependent, or covered by an SLA? |
| Independence | Is there a copy outside the primary database account/provider boundary? |
| Security | Encryption, immutable retention, least privilege, key separation, auditability? |
| Cost | Baseline platform cost vs incremental backup cost; storage, operations, egress, compute, standby, and drill costs? |
| Operations | Scheduling reliability, stale-backup alerts, credential rotation, and restore complexity? |
Use the arithmetic tool for cost scenarios. Show assumptions and distinguish total platform cost from incremental backup cost. Re-check vendor facts before every decision; pricing and retention change.
4. Prefer layered recovery
For small systems, a common low-cost shape is:
- Managed daily/continuous provider backup as the primary path
- Client-side-encrypted logical export to independent private object storage
- Immutable object names plus retention lock/lifecycle
- Automated freshness and integrity checks
- Regular isolated restore drills
The independent export is defense in depth, not permission to weaken the managed path. A third-party scheduler must not be the sole RPO control unless its timing and failure behavior are qualified.
5. Design encryption and authority boundaries
- Encrypt before upload when backups contain sensitive data.
- Keep the decryption key outside the repository, scheduler, storage provider, and backup payload.
- Ensure authorized responders can retrieve the key during an incident; untested offline custody creates an unusable backup.
- Use a dedicated least-privilege backup principal.
- Separate automated write credentials from restore/read and retention-policy administration.
- Never print connection strings, tokens, keys, or dump content.
- Delete plaintext and staging artifacts on success and failure.
- Store a manifest with migration head, ciphertext size, and cryptographic checksum or signature.
6. Define retention and immutability together
Retention is not just lifecycle deletion. Protect recovery points against overwrite and premature deletion with object lock, bucket lock, or equivalent. Align retention with privacy deletion propagation: document when deleted user data ages out and how emergency restores avoid silently resurrecting it.
7. Build a timed restore gate
Reserve part of RTO for incident authorization and application validation. For a one-hour RTO, a reasonable starting budget is 45 minutes for technical restore and 15 minutes for checks and reopening, but the owner must approve the split.
A drill passes only if it:
- Uses production-sized or explicitly bounded representative data
- Starts from an isolated clean target
- Replays the clean migration chain and exercises the upgrade path where applicable
- Verifies row counts, referential integrity, canonical checksums, migration versions, and security catalog
- Exercises runtime behavior under non-owner roles
- Restores or explicitly reconfigures auth, files, secrets, endpoints, and workers/functions
- Cleans temporary dumps and resources on success and failure
- Records stable machine-readable evidence without secrets
- Completes end to end within RTO
Requalify after material data growth, schema/security changes, provider changes, failed backups/drills, and at the approved recurring cadence.
8. Define escalation before the drill fails
Pre-approve the decision ladder:
- Optimize dump/restore and remove avoidable manual steps.
- Add a provisioned recovery target or warm standby.
- Move to PITR/continuous archiving or replication.
- Buy a provider SLA/support tier if contractual assurance is required.
Do not claim the RTO while the timed gate is red.
Automation and Alerts
Alert on:
- Job failure or missed schedule
- Backup age beyond policy
- Empty or unexpectedly small/large output
- Encryption or manifest failure
- Upload/read-back failure
- Integrity mismatch
- Retention-lock/lifecycle drift
- Restore-drill failure or duration regression
Test the alert path; a silent scheduled job is not an operational control.
Pitfalls
- Choosing numbers but no system: RPO/RTO values alone do not activate a recovery policy.
- Calling daily scheduling proof of 24-hour RPO: completion, freshness, integrity, and recoverability must be monitored.
- Equating provider restore with service recovery: configuration, secrets, files, auth, validation, and reopening consume RTO.
- Backing up only PostgreSQL: object files and external configuration may be omitted.
- Provider-managed encryption only: it may not satisfy key-separation requirements for sensitive exports.
- One credential can write, read, delete, and change retention: compromise destroys production and recovery evidence.
- Testing against tiny fixtures only: a size-dependent restore needs production-sized timing evidence.
- Publishing precise cost without assumptions: calculate retained GB-month, operations, runner time, compute, and standby separately.
- Selecting a provider before approval/drill: document it as a candidate and keep the production contract fail closed.
Verification Checklist
- RPO and RTO and their measurement boundaries are explicit.
- Every stateful data/configuration class has a recovery path or explicit exclusion.
- Current official docs and pricing support provider and cost claims.
- Baseline and incremental costs are separated with assumptions.
- Backup credentials and recovery keys follow least privilege and separation of duties.
- Independent copies are encrypted, immutable for the retention window, and lifecycle-managed.
- Freshness, integrity, policy drift, and restore failures alert.
- A production-sized isolated restore completed within RTO and produced evidence.
- Rollback versus roll-forward rules and incident authorization are documented.
- Privacy deletion propagation through backups is documented.
- The production policy remains inactive until every required decision and gate passes.
References
- For a dated, source-backed managed PostgreSQL plus object-storage cost pattern, see
references/managed-postgres-low-cost-pattern.md. Re-verify every vendor fact before reuse.
Supporting files: this skill's supporting files are held in the docsite at
docs/15-skills/_support/engineering/backup-and-disaster-recovery/— fetch them fresh fromjknash/docsitemain alongside this page. Source:jknash/hermes-shared-skills· branchhermes-jkdev001@1d0d545c3970·skills/engineering/backup-and-disaster-recovery/· view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.
version 1.0.0.
Published by Muse · 2026-10-04.