Disaster Recovery (DR) & Business Continuity Plan (BCP)
Proposed architecture and targets, not a production readiness certificate. The standby topology, WAL retention, backup schedule and numerical objectives below have not been verified against the deployed Railway environment. Actual recovery objectives must be configured and evidence independently reviewed in the admin recovery register. A disposable synthetic database restore does not establish hosted recovery time, data-loss exposure or external provider recovery.
1. Executive Summary & Recovery Objectives
As an e-commerce platform processing live financial transactions, handling campus logistics, and maintaining escrow liabilities, Debelu treats Business Continuity and Disaster Recovery as first-order engineering constraints.
The platform defines strict, measurable recovery targets categorized by operational criticality:
graph LR
subgraph Criticality Tiers
T1[Tier 1: Core Financial & Checkout Engine]
T2[Tier 2: Browse, Catalog & AI Services]
T3[Tier 3: Analytics & Back-Office Exports]
end
subgraph RTO / RPO Targets
T1 -->|RTO <= 1 Hour| RTO1[RPO <= 5 Minutes]
T2 -->|RTO <= 4 Hours| RTO2[RPO <= 1 Hour]
T3 -->|RTO <= 24 Hours| RTO3[RPO <= 24 Hours]
end
style T1 fill:#d4edda,stroke:#28a745,stroke-width:2px
style T2 fill:#fff3cd,stroke:#ffc107,stroke-width:2px
style T3 fill:#e2e3e5,stroke:#6c757d,stroke-width:1px1.1 Recovery Objectives Matrix
- Tier 1 (Checkout, Payment Ingestion, Escrow Ledger, Delivery PIN Verification):
- RTO (Recovery Time Objective): $\le 1\text{ Hour}$. Maximum duration from declared disaster to transaction restoration.
- RPO (Recovery Point Objective): $\le 5\text{ Minutes}$. Maximum acceptable data loss window via continuous Write-Ahead Log (WAL) archiving.
- Financial Double-Entry Invariant: $\text{RPO} = 0$ for confirmed payment webhooks and completed order handovers.
- Tier 2 (Storefront Catalog Browsing, Search, Reviews, Marketing App, Nduzi AI):
- $\text{RTO} \le 4\text{ Hours} \quad|\quad \text{RPO} \le 1\text{ Hour}$.
- Tier 3 (Bulk Audit Exports, Campaign Analytics, Historical Reporting):
- $\text{RTO} \le 24\text{ Hours} \quad|\quad \text{RPO} \le 24\text{ Hours}$.
2. Infrastructure Topology & Multi-Cloud Redundancy
Debelu operates on a decoupled multi-cloud architecture engineered to eliminate single points of failure (SPOF):
graph TD
subgraph Edge & Routing Layer
CF[Cloudflare Edge Network & Pages]
DNS[Cloudflare Anycast DNS]
WAF[Cloudflare WAF & DDoS Shield]
end
subgraph Compute Layer
FLY_PRI[Fly.io Primary: London LHR]
FLY_SEC[Fly.io Standby: Frankfurt FRA]
end
subgraph Primary Data Plane
SUPA_DB[(Supabase Managed Postgres)]
SUPA_WAL[Continuous WAL Stream Archive]
SUPA_STORE[Supabase Object Storage]
end
subgraph Independent Disaster Backup Vault
R2[(Cloudflare R2 Encrypted Cold Storage)]
end
subgraph External Financial Rail
PS[Paystack Payment Gateway]
end
DNS --> WAF --> CF
WAF --> FLY_PRI
FLY_PRI -.->|Failover| FLY_SEC
FLY_PRI --> SUPA_DB
FLY_SEC --> SUPA_DB
FLY_PRI --> PS
SUPA_DB --> SUPA_WAL
SUPA_DB -->|Nightly Logical pg_dump| R23. Database Disaster Recovery & Point-in-Time Recovery (PITR)
3.1 Continuous WAL Archiving (PITR)
- Supabase Postgres streams real-time Write-Ahead Logs (WAL) continuously to isolated secondary storage.
- Granular Rollback: Allows restoration of the database to any specified second within a rolling 7-day retention window.
- Corruption Recovery: If an accidental migration, erroneous script, or software bug corrupts data at 14:32:15, on-call engineers initiate recovery to 14:32:10, losing less than 5 seconds of non-financial transactions.
3.2 Offsite Encrypted Logical Backups (Cloudflare R2 Vault)
- In addition to hosted platform backups, an automated nightly Cron job executes
pg_dumptargeting an independent provider:bash# Nightly encrypted dump pipeline pg_dump --clean --if-exists --no-owner --no-privileges "$DATABASE_URL" \ | gpg --symmetric --cipher-algo AES256 --batch --passphrase "$BACKUP_PASSPHRASE" \ | aws s3 cp - "s3://debelu-dr-vault/backups/$(date +%Y-%m-%d).sql.gpg" \ --endpoint-url "https://$CLOUDFLARE_R2_ACCOUNT_ID.r2.cloudflarestorage.com" - Automated lifecycle policy: Retained for 30 days before automated deletion.
4. Disaster Scenarios & Failover Playbooks
Scenario A: Supabase Regional Data Center Outage
Trigger: Supabase AWS region experiences unrecoverable physical failure or extended network partition ($> 15$ mins).
- Declare SEV-1: Incident Commander convenes the incident response channel.
- Engage Maintenance Edge:bash
# Route traffic to Cloudflare Maintenance Page via Worker wrangler kv:key put --binding=PLATFORM_STATE "global_maintenance" "true" - Spin Up Standby Postgres: Provision an emergency managed PostgreSQL 16 cluster on Neon or AWS RDS.
- Restore Latest Snapshot: Decrypt and restore the latest R2 logical dump.
- Update Compute Secrets:bash
fly secrets set DATABASE_URL="postgres://user:pass@standby-host:5432/debelu" - Execute Migration Checks: Run
npm run check:driftto confirm zero schema variance. - Lift Maintenance: Route live traffic back to storefront.
Scenario B: Paystack Payment Gateway Outage
Trigger: Paystack API returns 5xx errors or timeouts exceed 15s across 5 consecutive transactions.
- Automated Circuit Breaker: The storefront payment service catches recurring initialization timeouts.
- Graceful Degradation:
- Card and USSD checkouts temporarily display: "Payment Gateway Experiencing Delays".
- Campus orders transition automatically to "Pay on Pickup / Campus Hub Reserve" mode.
- Buyer checkout intent is held in
reservedstate without charging the customer.
- Webhook Replay Recovery: When Paystack recovers, the Edge Function replays queued webhook events from the dead-letter queue.
Scenario C: Fly.io Regional Compute Crash
Trigger: London (LHR) Fly.io gateway experiences host hardware failures.
- Anycast Automatic Routing: Fly.io Anycast IPs automatically reroute incoming TCP connections to the standby instances running in Frankfurt (FRA).
- Machine Scaling: Standby machines scale up from min-replicas ($2$) to full capacity ($10$) within 90 seconds.
5. Automated Recovery Drills (command-center-recovery-drill.mjs)
Debelu maintains a self-contained, in-memory recovery verification runner ([scripts/command-center-recovery-drill.mjs](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/scripts/command-center-recovery-drill.mjs)) that executes 8 representative financial and governance disaster fixtures:
# Execute isolated disaster recovery drill in local runtime
node scripts/command-center-recovery-drill.mjs --output .artifacts/recovery-drill.jsonThe 8 Verification Fixtures
checkout-reservation: Verifies payment intent locking, timeout uncertainty, and race-condition defenses.payout-dispatch: Verifies transfer claim tokens, rate limits, and double-withdrawal prevention.payout-reconciliation: Verifies Maker-Checker dual authorization and Paystack outcome drift rejection.private-export: Verifies cryptographic SHA-256 export packaging and AAL2 step-up access.campaign-publication: Verifies marketing campaign publication boundaries.wallet-refund: Verifies atomic wallet balance refund and immutable transaction receipts.scoped-privacy-erasure: Verifies the 7 erasable collections vs 7 statutory retained sources.expanded-private-export: Verifies multi-table entity discovery and schema compliance.
Cryptographic Receipt Verification
The runner outputs an immutable evidence receipt (debelu_isolated_recovery_drill_v1) containing SHA-256 digests of all migration files, Git commit hashes, and test results:
# Verify integrity of an existing recovery drill receipt
node scripts/command-center-recovery-drill.mjs --verify .artifacts/recovery-drill.json6. Business Continuity Drill Cadence & Audit SLAs
| Activity | Frequency | Target Objective | Participants |
|---|---|---|---|
| Automated Recovery Fixture Run | Every Git Commit (CI) | Confirm zero regression across the 8 critical financial and privacy drill fixtures. | GitHub Actions CI |
| Tabletop Disaster Simulation | Quarterly | Simulated scenario walk-through (e.g. AWS eu-central total blackout, DNS hijack). | Engineering & Ops Leads |
| Live Database Restoration Drill | Bi-Annual | Restoring encrypted R2 logical dump into an isolated staging environment and running smoke suites. | DevOps / DBA Lead |
| Blameless Post-Mortem Audit | Post-Incident | Mandatory root-cause analysis (RCA) published within 48 hours of any SEV-1 or SEV-2 event. | Incident Commander & Team |