Skip to content

Disaster Recovery (DR) & Business Continuity Plan (BCP) ​

Proposed architecture and targets, not a production readiness certificate. The standby topology, WAL retention, backup schedule and numerical objectives below have not been verified against the deployed Railway environment. Actual recovery objectives must be configured and evidence independently reviewed in the admin recovery register. A disposable synthetic database restore does not establish hosted recovery time, data-loss exposure or external provider recovery.


1. Executive Summary & Recovery Objectives ​

As an e-commerce platform processing live financial transactions, handling campus logistics, and maintaining escrow liabilities, Debelu treats Business Continuity and Disaster Recovery as first-order engineering constraints.

The platform defines strict, measurable recovery targets categorized by operational criticality:

mermaid
graph LR
    subgraph Criticality Tiers
        T1[Tier 1: Core Financial & Checkout Engine]
        T2[Tier 2: Browse, Catalog & AI Services]
        T3[Tier 3: Analytics & Back-Office Exports]
    end

    subgraph RTO / RPO Targets
        T1 -->|RTO <= 1 Hour| RTO1[RPO <= 5 Minutes]
        T2 -->|RTO <= 4 Hours| RTO2[RPO <= 1 Hour]
        T3 -->|RTO <= 24 Hours| RTO3[RPO <= 24 Hours]
    end

    style T1 fill:#d4edda,stroke:#28a745,stroke-width:2px
    style T2 fill:#fff3cd,stroke:#ffc107,stroke-width:2px
    style T3 fill:#e2e3e5,stroke:#6c757d,stroke-width:1px

1.1 Recovery Objectives Matrix ​

  • Tier 1 (Checkout, Payment Ingestion, Escrow Ledger, Delivery PIN Verification):
    • RTO (Recovery Time Objective): $\le 1\text{ Hour}$. Maximum duration from declared disaster to transaction restoration.
    • RPO (Recovery Point Objective): $\le 5\text{ Minutes}$. Maximum acceptable data loss window via continuous Write-Ahead Log (WAL) archiving.
    • Financial Double-Entry Invariant: $\text{RPO} = 0$ for confirmed payment webhooks and completed order handovers.
  • Tier 2 (Storefront Catalog Browsing, Search, Reviews, Marketing App, Nduzi AI):
    • $\text{RTO} \le 4\text{ Hours} \quad|\quad \text{RPO} \le 1\text{ Hour}$.
  • Tier 3 (Bulk Audit Exports, Campaign Analytics, Historical Reporting):
    • $\text{RTO} \le 24\text{ Hours} \quad|\quad \text{RPO} \le 24\text{ Hours}$.

2. Infrastructure Topology & Multi-Cloud Redundancy ​

Debelu operates on a decoupled multi-cloud architecture engineered to eliminate single points of failure (SPOF):

mermaid
graph TD
    subgraph Edge & Routing Layer
        CF[Cloudflare Edge Network & Pages]
        DNS[Cloudflare Anycast DNS]
        WAF[Cloudflare WAF & DDoS Shield]
    end

    subgraph Compute Layer
        FLY_PRI[Fly.io Primary: London LHR]
        FLY_SEC[Fly.io Standby: Frankfurt FRA]
    end

    subgraph Primary Data Plane
        SUPA_DB[(Supabase Managed Postgres)]
        SUPA_WAL[Continuous WAL Stream Archive]
        SUPA_STORE[Supabase Object Storage]
    end

    subgraph Independent Disaster Backup Vault
        R2[(Cloudflare R2 Encrypted Cold Storage)]
    end

    subgraph External Financial Rail
        PS[Paystack Payment Gateway]
    end

    DNS --> WAF --> CF
    WAF --> FLY_PRI
    FLY_PRI -.->|Failover| FLY_SEC
    FLY_PRI --> SUPA_DB
    FLY_SEC --> SUPA_DB
    FLY_PRI --> PS

    SUPA_DB --> SUPA_WAL
    SUPA_DB -->|Nightly Logical pg_dump| R2

3. Database Disaster Recovery & Point-in-Time Recovery (PITR) ​

3.1 Continuous WAL Archiving (PITR) ​

  • Supabase Postgres streams real-time Write-Ahead Logs (WAL) continuously to isolated secondary storage.
  • Granular Rollback: Allows restoration of the database to any specified second within a rolling 7-day retention window.
  • Corruption Recovery: If an accidental migration, erroneous script, or software bug corrupts data at 14:32:15, on-call engineers initiate recovery to 14:32:10, losing less than 5 seconds of non-financial transactions.

3.2 Offsite Encrypted Logical Backups (Cloudflare R2 Vault) ​

  • In addition to hosted platform backups, an automated nightly Cron job executes pg_dump targeting an independent provider:
    bash
    # Nightly encrypted dump pipeline
    pg_dump --clean --if-exists --no-owner --no-privileges "$DATABASE_URL" \
      | gpg --symmetric --cipher-algo AES256 --batch --passphrase "$BACKUP_PASSPHRASE" \
      | aws s3 cp - "s3://debelu-dr-vault/backups/$(date +%Y-%m-%d).sql.gpg" \
        --endpoint-url "https://$CLOUDFLARE_R2_ACCOUNT_ID.r2.cloudflarestorage.com"
  • Automated lifecycle policy: Retained for 30 days before automated deletion.

4. Disaster Scenarios & Failover Playbooks ​

Scenario A: Supabase Regional Data Center Outage ​

Trigger: Supabase AWS region experiences unrecoverable physical failure or extended network partition ($> 15$ mins).

  1. Declare SEV-1: Incident Commander convenes the incident response channel.
  2. Engage Maintenance Edge:
    bash
    # Route traffic to Cloudflare Maintenance Page via Worker
    wrangler kv:key put --binding=PLATFORM_STATE "global_maintenance" "true"
  3. Spin Up Standby Postgres: Provision an emergency managed PostgreSQL 16 cluster on Neon or AWS RDS.
  4. Restore Latest Snapshot: Decrypt and restore the latest R2 logical dump.
  5. Update Compute Secrets:
    bash
    fly secrets set DATABASE_URL="postgres://user:pass@standby-host:5432/debelu"
  6. Execute Migration Checks: Run npm run check:drift to confirm zero schema variance.
  7. Lift Maintenance: Route live traffic back to storefront.

Scenario B: Paystack Payment Gateway Outage ​

Trigger: Paystack API returns 5xx errors or timeouts exceed 15s across 5 consecutive transactions.

  1. Automated Circuit Breaker: The storefront payment service catches recurring initialization timeouts.
  2. Graceful Degradation:
    • Card and USSD checkouts temporarily display: "Payment Gateway Experiencing Delays".
    • Campus orders transition automatically to "Pay on Pickup / Campus Hub Reserve" mode.
    • Buyer checkout intent is held in reserved state without charging the customer.
  3. Webhook Replay Recovery: When Paystack recovers, the Edge Function replays queued webhook events from the dead-letter queue.

Scenario C: Fly.io Regional Compute Crash ​

Trigger: London (LHR) Fly.io gateway experiences host hardware failures.

  1. Anycast Automatic Routing: Fly.io Anycast IPs automatically reroute incoming TCP connections to the standby instances running in Frankfurt (FRA).
  2. Machine Scaling: Standby machines scale up from min-replicas ($2$) to full capacity ($10$) within 90 seconds.

5. Automated Recovery Drills (command-center-recovery-drill.mjs) ​

Debelu maintains a self-contained, in-memory recovery verification runner ([scripts/command-center-recovery-drill.mjs](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/scripts/command-center-recovery-drill.mjs)) that executes 8 representative financial and governance disaster fixtures:

bash
# Execute isolated disaster recovery drill in local runtime
node scripts/command-center-recovery-drill.mjs --output .artifacts/recovery-drill.json

The 8 Verification Fixtures ​

  1. checkout-reservation: Verifies payment intent locking, timeout uncertainty, and race-condition defenses.
  2. payout-dispatch: Verifies transfer claim tokens, rate limits, and double-withdrawal prevention.
  3. payout-reconciliation: Verifies Maker-Checker dual authorization and Paystack outcome drift rejection.
  4. private-export: Verifies cryptographic SHA-256 export packaging and AAL2 step-up access.
  5. campaign-publication: Verifies marketing campaign publication boundaries.
  6. wallet-refund: Verifies atomic wallet balance refund and immutable transaction receipts.
  7. scoped-privacy-erasure: Verifies the 7 erasable collections vs 7 statutory retained sources.
  8. expanded-private-export: Verifies multi-table entity discovery and schema compliance.

Cryptographic Receipt Verification ​

The runner outputs an immutable evidence receipt (debelu_isolated_recovery_drill_v1) containing SHA-256 digests of all migration files, Git commit hashes, and test results:

bash
# Verify integrity of an existing recovery drill receipt
node scripts/command-center-recovery-drill.mjs --verify .artifacts/recovery-drill.json

6. Business Continuity Drill Cadence & Audit SLAs ​

ActivityFrequencyTarget ObjectiveParticipants
Automated Recovery Fixture RunEvery Git Commit (CI)Confirm zero regression across the 8 critical financial and privacy drill fixtures.GitHub Actions CI
Tabletop Disaster SimulationQuarterlySimulated scenario walk-through (e.g. AWS eu-central total blackout, DNS hijack).Engineering & Ops Leads
Live Database Restoration DrillBi-AnnualRestoring encrypted R2 logical dump into an isolated staging environment and running smoke suites.DevOps / DBA Lead
Blameless Post-Mortem AuditPost-IncidentMandatory root-cause analysis (RCA) published within 48 hours of any SEV-1 or SEV-2 event.Incident Commander & Team

Released under Proprietary Enterprise License.