Skip to content

Incident Response Playbook & Disaster Escalation ​


1. Executive Summary & Incident Command System (ICS) ​

Debelu's Incident Response framework provides a deterministic, military-grade operational protocol for identifying, triaging, containing, mitigating, and reviewing production disruptions across all five production surfaces (Storefront, Marketing, Backend API, Edge Functions, and Supabase Database).

The architecture adheres to the Federal Emergency Management Agency (FEMA) Incident Command System (ICS) adapted for Tier-1 technology organizations. During any SEV-1 or SEV-2 incident, normal organizational hierarchy is suspended; operational authority vests entirely in the Incident Commander (IC).

mermaid
graph TD
    subgraph Incident Command Structure
        IC[Incident Commander: Overall Incident Authority]
        TL[Technical Lead: Operations & Code Remediation]
        CL[Communications Lead: Status Page & Stakeholder Messaging]
        SC[Scribe: Timeline Logging & War Room Ledger]
    end

    subgraph Remediation Squads
        TL --> DB_ENG[Database & Infrastructure Eng]
        TL --> FIN_ENG[Payments & Ledger Eng]
        TL --> SEC_ENG[Security & Privacy Eng]
    end

    subgraph Communication Channels
        CL --> STATUS[debelu.com/status Public Updates]
        CL --> EXEC[Executive Briefings & Slack #incidents]
        CL --> SUPP[Customer Support Canned Responses]
    end

2. Severity Classification Matrix ​

Severity is evaluated immediately upon detection based on blast radius, user impact, financial risk, and data integrity:

SeverityDefinition & Technical ImpactTarget TriageTarget MTTRNotification CadenceIncident Commander Assignment
SEV-1: CriticalComplete service outage; payment capture failure; corrupted double-entry ledger; security compromise / data breach.$\le 5\text{ mins}$$\le 45\text{ mins}$Every $15\text{ mins}$VP of Engineering / Head of Tech
SEV-2: MajorPartial service outage; degraded checkout flow; campus order dispatch stalled; high API error rates ($> 5%$).$\le 15\text{ mins}$$\le 2\text{ hours}$Every $30\text{ mins}$Staff Engineer / Tech Lead
SEV-3: MinorNon-critical feature broken (e.g. Nduzi AI responses delayed, reviews not rendering); single-campus UI defect.$\le 1\text{ hour}$$\le 8\text{ hours}$Daily at StandupSenior On-Call Engineer
SEV-4: LowCosmetic rendering bugs; minor admin table pagination latency; internal documentation gaps.$\le 4\text{ hours}$Next SprintSprint TriageAssigned Feature Engineer

3. Incident Command Roles & Responsibilities ​

                                  INCIDENT COMMAND CHAIN
                                            │
               ┌────────────────────────────┼────────────────────────────┐
               ▼                            ▼                            ▼
      Incident Commander              Technical Lead            Communications Lead
     - Ultimate authority          - Directs debug plane       - Owns status page
     - Coordinates squads          - Authorizes code rollbacks - Updates exec team
     - Enforces timeboxes          - Deploys hotfixes          - Briefs support

3.1 Incident Commander (IC) ​

  • Single Source of Truth: The IC does not write code or debug logs directly. Their singular duty is to maintain situational awareness, allocate engineering squads, enforce timeboxes (e.g., "If rollback fails in 10 minutes, we fail over to read-only replica"), and prevent chaotic thrashing.
  • Command Handover Protocol: If an incident spans longer than 4 hours, formal handover must occur via Slack voice bridge:
    1. Current state summary and active mitigation hypotheses.
    2. Inventory of changes deployed within the last 60 minutes.
    3. Formal statement: "I, [Name A], hereby transfer Incident Command to [Name B]."

3.2 Technical Lead (TL) ​

  • Commands the debugging plane. Assigns engineers to isolate root cause, evaluates database transaction locks (pg_stat_activity), reviews Sentry breadcrumbs, and drafts hotfix PRs or rollback commands.

3.3 Communications Lead (CL) ​

  • Owns public and internal messaging. Shields the IC and TL from executive or customer interruptions.
  • Updates [debelu.com/status](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-marketing/app/status/page.tsx) and dispatches customer care talk tracks.

3.4 Scribe ​

  • Maintains the real-time chronological ledger in the designated incident document: timestamps, tool outputs, hypotheses tested, configuration values changed, and rollback tags.

4. Communication Protocols & War Room Orchestration ​

4.1 War Room Infrastructure ​

Immediately upon SEV-1/SEV-2 declaration:

  1. Slack Channel: Automated creation of #incident-YYYYMMDD-<title> via Sentry/PagerDuty integration.
  2. War Room Video Bridge: Dedicated Google Meet opened and pinned to the Slack topic.
  3. Executive Broadcast: Summary posted to #incidents-announcements:

    [SEV-1 ACTIVE] Paystack Webhook Ingestion Timeout
    Impact: Checkout confirmations delayed; order state transitions stalled.
    IC: @chisom | TL: @frankline | War Room: meet.google.com/deb-inci-war

4.2 Public Status Page Messaging Guidelines ​

Updates to the public status page must be concise, objective, and devoid of internal blame or technical jargon:

  • Investigating: "We are currently investigating reports of delays in order checkout and payment confirmation. Our engineering team is actively diagnosing the issue. Next update in 15 minutes."
  • Identified: "We have identified an upstream connectivity disruption affecting payment gateway handshakes. Remediation is underway."
  • Monitoring: "A fix has been deployed and payment confirmations are recovering. We are monitoring transaction queues to ensure backlog clearance."
  • Resolved: "All payment services have fully recovered. Any unconfirmed orders have been automatically reconciled with bank ledgers."

5. Automated Operational Playbooks ​

mermaid
flowchart TD
    ALARM{Incoming Alarm}
    ALARM -->|Payment Spikes| PB_PAY[Playbook 1: Payment Gateways]
    ALARM -->|HealthCheck Failure| PB_DB[Playbook 2: Database Outage]
    ALARM -->|502 Bad Gateway| PB_API[Playbook 3: Backend API Crash]
    ALARM -->|Breach / Token Leak| PB_SEC[Playbook 4: Security Breach]
    ALARM -->|BullMQ Stalled| PB_QUEUE[Playbook 5: Queue Backlog]

5.1 Playbook 1: Payment Processing & Webhook Failure ​

Symptoms ​

  • Sentry alert: PaymentService.initializeTransaction failing with timeout or HTTP 502.
  • Webhook ingestion queue depth rising; buyer wallets debited without order creation.

Step-by-Step Triage ​

  1. Verify Paystack Gateway Health: Check upstream status: curl -I https://status.paystack.com/ and Paystack API ping.
  2. Inspect Backend Webhook Ingestion:
    bash
    # Check logs for raw body buffer verification errors or timeouts
    railway logs --service debelu-backend | grep -i "paystack-webhook"
  3. Execute Emergency Webhook Queue Catch-Up: If webhooks failed during an API blip, execute the historical reconciliation daemon:
    bash
    node scripts/reconcile-payouts.mjs --dry-run=false --window=24h
  4. Trigger Double-Entry Restitution Check: Query unpaid order intents with successful Paystack charge IDs:
    sql
    SELECT id, order_id, amount_minor, provider_reference 
    FROM public.checkout_payment_intents 
    WHERE status = 'requires_action' 
      AND updated_at < NOW() - INTERVAL '15 minutes';
  5. Mitigation: If Paystack is hard down, enable the platform payment circuit breaker in Command Center:
    typescript
    // Sets platform_controls.maintenance_mode or disables card checkout
    await platformControlsService.update({ checkout_enabled: false });

5.2 Playbook 2: Database Outage & Connection Exhaustion ​

Symptoms ​

  • HealthCheckService.checkDatabase() returns unhealthy.
  • PostgreSQL error 53300 (too_many_connections) or 57014 (query_canceled).
  • High latency spikes on API endpoints.

Step-by-Step Triage ​

  1. Identify Blocking Locks & Runaway Queries: Connect via Supabase Direct SQL console:
    sql
    SELECT pid, now() - pg_stat_activity.query_start AS duration, query, state
    FROM pg_stat_activity
    WHERE (now() - pg_stat_activity.query_start) > interval '10 seconds'
      AND state != 'idle';
  2. Terminate Runaway Queries:
    sql
    SELECT pg_terminate_backend(pid) 
    FROM pg_stat_activity 
    WHERE duration > interval '30 seconds' AND pid <> pg_backend_pid();
  3. Activate Read-Only Circuit Breakers: If connection pool is saturated, trip write circuit breakers in [circuitBreakers.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/middleware/circuitBreakers.ts) to protect core catalog browsing while halting heavy analytics queries.
  4. Point-In-Time Recovery (PITR): If data corruption occurred, follow [docs/operations/disaster-recovery-and-bcp.md](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/docs/operations/disaster-recovery-and-bcp.md) to initiate Supabase WAL recovery to $T_{-1\text{min}}$.

5.3 Playbook 3: Backend API Crash Loops & OOM ​

Symptoms ​

  • Railway service restarting repeatedly with exit code 137 (SIGKILL / Out of Memory).
  • Storefront returning 502 Bad Gateway or 504 Gateway Timeout.

Step-by-Step Triage ​

  1. Inspect Memory Telemetry & Stack Trace:
    bash
    railway logs --service debelu-backend -n 100
  2. Immediate Fast Rollback: If caused by a recent deployment:
    bash
    railway rollback --service debelu-backend
  3. Memory Profile Dump: If OOM persists on stable release, check BullMQ queue memory footprint and verify that heavy memory jobs (e.g. bulk CSV export, image processing) are sandboxed to background workers.
  4. Scale Container Resources Temporarily: Upgrade Railway container from 1 GB to 4 GB RAM to absorb memory spikes during triage.

5.4 Playbook 4: Security Breach & Unauthorized Access ​

Symptoms ​

  • Suspected leak of SUPABASE_SERVICE_ROLE_KEY or PAYSTACK_SECRET_KEY.
  • High-volume unauthorized data extraction from public.users or public.vendors.
  • Brute-force attacks bypassing rate limits.

Step-by-Step Triage ​

  1. Isolate Compromised Credentials: Immediately follow [docs/operations/runbooks/secrets-rotation.md](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/docs/operations/runbooks/secrets-rotation.md) to revoke and rotate the affected secret.
  2. Block Offending IP Addresses: Add malicious IP subnets to Cloudflare WAF block list.
  3. Invalidate Active Staff Sessions:
    sql
    -- Force immediate logout of all staff sessions
    UPDATE auth.users 
    SET raw_app_meta_data = raw_app_meta_data || '{"revoked_at":"' || NOW() || '"}'::jsonb 
    WHERE raw_user_meta_data->>'role' IN ('admin', 'superadmin', 'support_agent');
  4. Statutory Regulatory Clock (NDPC Notification): Under the Nigeria Data Protection Act 2023 (NDPA), personal data breaches impacting Nigerian data subjects must be reported to the Nigeria Data Protection Commission (NDPC) within 72 hours of confirmation.

5.5 Playbook 5: BullMQ Queue Backlog & Worker Stall ​

Symptoms ​

  • QueueObservationsService reports queue depth $> 5,000$ jobs.
  • Vendor notifications or AI chat streaming jobs delayed.

Step-by-Step Triage ​

  1. Inspect Worker Heartbeats:
    bash
    node scripts/check-command-center-health.mjs
  2. Inspect Dead-Letter Queue (DLQ): Check stalled jobs in Redis:
    bash
    node -e '
      const { Queue } = require("bullmq");
      const q = new Queue("notifications", { connection: { host: process.env.REDIS_HOST } });
      q.getFailed().then(jobs => console.log("Failed jobs:", jobs.length));
    '
  3. Drain Corrupt Poison-Pill Jobs: If a single malformed job payload is crashing worker loops:
    bash
    node scripts/drain-poison-pill.mjs --queue=notifications --jobId=<JOB_ID>

6. Recovery Drills & Automated Chaos Simulation ​

Debelu enforces scheduled quarterly disaster recovery drills to validate operational readiness and runbook accuracy:

6.1 Automated Recovery Test Suite ​

Repository verification tools:

  • [scripts/command-center-recovery-drill.mjs](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/scripts/command-center-recovery-drill.mjs): Simulates total Command Center degradation and executes step-by-step restoration of system controls.
  • [scripts/recovery-drill.test.mjs](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/scripts/recovery-drill.test.mjs): Vitest/Jest suite validating database fallback and cache failover under degraded network states.

6.2 Drill Execution Protocol ​

bash
# Execute dry-run disaster drill in staging
npm run drill:recovery -- --environment=staging

Every completed drill requires a signed scorecard documenting:

  • Actual detection time vs alerting threshold.
  • Total elapsed time to execute failover commands.
  • Data loss verification (RPO audit confirming $0$ dropped financial transactions).

7. Post-Incident Review (PIR) & Blameless Governance ​

A Post-Incident Review is mandatory for all SEV-1 and SEV-2 incidents, to be completed within 72 hours of incident resolution.

7.1 The Blameless Post-Mortem Philosophy ​

  • We do not punish human error; we improve systemic resilience.
  • Assumption: Engineers make the best possible decisions given the information, tooling, and constraints available to them at the time.
  • Goal: Uncover systemic design flaws, missing alerting thresholds, fragile dependencies, and incomplete documentation.

7.2 Post-Mortem Document Template ​

The Incident Commander is responsible for publishing the PIR artifact to docs/incidents/YYYY-MM-DD-<title>.md:

markdown
# Post-Incident Review: [Incident Title]

**Date**: YYYY-MM-DD  
**Severity**: SEV-1 | SEV-2  
**Incident Commander**: [Name]  
**Technical Lead**: [Name]  
**Time to Detect (TTD)**: XX mins  
**Time to Mitigate (TTM)**: XX mins  

## 1. Executive Summary
Brief non-technical overview of what happened, user impact, and resolution.

## 2. Customer Impact & Business Loss
- Number of affected users / orders:
- Total GMV delayed or at risk:
- SLA penalty impact:

## 3. Incident Timeline (UTC +1)
- **HH:MM** - Sentry alert fires for ...
- **HH:MM** - IC declares SEV-1; war room opened.
- **HH:MM** - Technical Lead identifies root cause in ...
- **HH:MM** - Rollback command issued via Railway CLI.
- **HH:MM** - Recovery verified; status page updated to Resolved.

## 4. Root Cause Analysis (5 Whys)
1. Why did the API crash? (Uncaught exception in webhook handler)
2. Why was it uncaught? (Missing schema validation on new provider field)
3. Why was validation missing? (Upstream API added an undocumented payload)
4. Why was this not caught in staging? (Staging mock webhook fixtures were outdated)
5. Why were mocks outdated? (No automated drift test between Paystack API and our schemas)

## 5. Corrective and Preventive Actions (CAPA)
| Action Item | Type (Prevent / Detect / Mitigate) | Owner | Target Date | JIRA / Ticket |
| :--- | :--- | :--- | :--- | :--- |
| Add Zod schema passthrough parsing for webhooks | Prevent | @chisom | 2026-10-12 | DEB-4091 |
| Create automated daily webhook contract test | Detect | @frankline | 2026-10-15 | DEB-4092 |
| Add fast-abort circuit breaker on JSON parse failures | Mitigate | @dev | 2026-10-18 | DEB-4093 |

Released under Proprietary Enterprise License.