Purpose
This SOP outlines the step-by-step process Tier-1 agents must follow when identifying, handling, and communicating outages. The goal is to ensure fast coordination, accurate updates, and minimal client escalation.
1. Identifying a Potential Outage
Tier-1 often notices outages first. Pay attention to these signals:
Tickets: Incoming P0/P1 tickets reporting:
- Endpoint down
- Messages not flowing
- 500 errors
Slack Channels:
- #mb-outage — Monitor alerts, updates, and discussions from internally identified outages.
- #asksupport — If an internal user raises an issue, clarify quickly. If it appears systemic, move the discussion to #mb-outage.
PagerDuty Alerts:
- Never ignore PagerDuty alerts.
- Treat every alert as a potential outage until confirmed.
- Quickly determine whether the issue is isolated or widespread.
Guideline: When unsure, begin investigation. Delays increase escalations.
2. Outage Classification: Full vs. Partial / Degraded
Before starting the bridge, quickly assess the scope. This determines urgency and communication tone.
| Type | Definition | Examples |
|---|---|---|
| Full Outage | A service is completely unavailable for all or most clients. | SMS delivery completely stopped, API returning 500 for all requests. |
| Partial Outage | A service is unavailable for a subset of clients or regions. | SMS delayed in US-East only, one client's API calls failing while others are fine. |
| Service Degradation | Service is functional but significantly slower or unreliable. | Message delivery taking 10x longer than normal, intermittent 503 errors. |
Why this matters: Full outages require immediate Status Page updates and leadership escalation. Partial outages and degradations may allow a longer investigation window before public communication, but still require a bridge and internal updates.
When in doubt, treat it as a full outage until confirmed otherwise.
3. Starting the Bridge (Must Happen Within 3–5 Minutes)
Once you sense an outage:
Step 1 — Start a Bridge Call within 5 minutes
-
Create the bridge immediately. Do not wait for confirmation.
- Share the bridge link in #mb-outage
- Pin the message so everyone can join quickly.
Step 2 — Ensure the Right Teams Are Present
Leadership: When an outage, service degradation, or escalation is detected and a bridge is being initiated, trigger an incident using the PagerDuty Alert Trigger Guide for CloudOps and Leadership Teams. This will alert available leadership members to respond and join the bridge.
-
CloudOps: Always add CloudOps via #mb-netops.
- Trigger PagerDuty: PagerDuty Alert Trigger Guide for CloudOps and Leadership Teams
- Backup method: Cloud Operations Contact List
- For escalation contacts: SOP - CO-Support - Critical Issue Process.docx
-
Tier-2: Always add Tier-2.
- Trigger PagerDuty: PagerDuty Tier-2
-
Client-Specific Issues: Use dedicated client channels (e.g., #client-ameren for Ameren).
- Full client/product contacts: Dev Contact Matrix – Convey
-
ECS / Eons2 Impact:
- Use #ecs-discussion to bring in the ECS team.
- Use #mb-eons-2 to bring in the Eons 2 team.
4. Escalation Scenarios & Decision Trees
Use the decision flows below to determine the correct escalation path quickly.
Decision Tree A — Do I Need to Start a Bridge?
Are you seeing P0/P1 tickets, PagerDuty alerts, or #mb-outage activity? │ ├── YES → Is it isolated to a single client or account? │ ├── YES → Investigate first (15 min). If it's outage, start bridge. │ └── NO → Start bridge immediately. Notify CloudOps + Tier-2. │ └── NO → Continue monitoring.Monitor #mb-outage proactively.
Decision Tree B — Who Do I Escalate To?
What type of issue is it?
│
├── Infrastructure / Network (endpoints down, DNS, routing)
│ └── → Escalate to CloudOps via #mb-netops + PagerDuty
│
├── Application / Platform (API errors, service logic, code issues)
│ └── → Escalate to Tier-2 via PagerDuty
│
├── Client-Specific (one account affected, others fine)
│ └── → Escalate to TAM for that client
│ → Use client-dedicated Slack channel
│ → Use Dev matrix to pull Dev Resource
│
├── ECS / Eons2 platform
│ ├── ECS → #ecs-discussion
│ └── Eons2 → #mb-eons-2
│
└── Unknown / Multiple systems
└── → Start bridge, notify CloudOps + Tier-2 + Leadership simultaneously
Decision Tree C — When Do I Update the Status Page?
Has 15 minutes passed since the bridge started?
│
├── NO → Continue investigation. Post internal update in #mb-outage only.
│
└── YES → Has the outage been confirmed by CloudOps or owning team?
├── YES → Trip the Status Page. Begin external communication.
└── NO → Extend investigation by 10 minutes.
If still unconfirmed → escalate to leadership for direction.
Decision Tree D — Is This a Critical Account?
Is the affected client SCE, PG&E, or Duke?
│
├── YES → Is the outage longer than 15 minutes?
│ ├── YES → Immediately loop in CSM.
│ │ TAM handle direct client comms in parallel.
│ └── NO → Monitor. If not resolved in 15 min, loop in TAM.
│
└── NO → Follow standard outage process.
Escalate to CSM only if client starts escalating directly.
5. Notification Protocol (Critical Step)
As soon as the bridge is active:
- Start a 15-minute investigation timer.
- Verify any recent deployments in the #deployments channel.
- Trigger all PagerDuty alerts: CloudOps on-call, Leadership pager, and Tier-2.
- At the end of 15 minutes, if the outage is confirmed:
- Trip the Public Status Page after getting confirmation from the relevant team on the bridge.
- Begin external communication.
6. Confirming Client & Service Impact (During Bridge)
Tier-1 drives the initial scoping discussion. Ask product/service owners the following:
| Question | Examples |
|---|---|
| Which services are affected? | SMS / Voice / Email / Push / API / UI |
| Which regions? | US-East, US-West, Canada, Global |
| Which clients? | Ameren, Duke, PG&E, etc. |
| Is there a recent deployment or change? | Check #deployments |
| Is the issue getting worse, stable, or improving? | Trend direction matters for comms |
Record the impact concisely.
Example: "SMS delivery delays in US-East affecting Ameren and Duke customers. Issue stable, CloudOps investigating."
Do not assume impact — confirm with the owning team before communicating externally.
7. Communication & Updates
Tier-1 is responsible for timely and consistent updates.
Internal Updates
- Post in #mb-outage every 30 minutes and start documenting notes in the Teams bridge for later RCA.
- Keep updates short, factual, and confirmation based.
- Always include current status, who is investigating, and next update time.
Example internal update:
Bridge active. CloudOps investigating. Current impact: SMS delays in US-East. No new deployments found. Issue appears infrastructure-related. Next update in 30 minutes or sooner.
External Updates (Status Page)
- Use the Status Page SOP: SOP-CO- Public Status Page – All Documents
- Create a client ticket if required — grab contacts from the MB Zendesk Implementation Tracker.
- Always follow client-specific status page procedures where available.
- Only post confirmed information — never assumptions or guesses.
- Update the Status Page whenever there is a meaningful change in status, not just on the 30-minute cycle.
8. Handling Critical Accounts
If the outage affects SCE, PG&E, or Duke:
- Immediately involve their CSMs and TAM if the outage is prolonged.
Refer to the TAM Guide: Managed Services Team
Refer to the Transition Status – CSM Wise.xlsx for CSM names. - TAM will handle direct client conversations in parallel with status page updates.
- Tier-1 continues owning internal coordination — do not hand off bridge ownership to the TAM.
Tip: Looping TAM/CSMs in early avoids unnecessary escalations to leadership.
9. Shift Handover During a Long Outage
If an outage spans across a shift change, a proper handover is critical. An incomplete handover is one of the most common causes of duplicate actions, missed updates, and client escalations.
When to Initiate Handover
- Any active outage or bridge still running at shift change time.
- Any unresolved P0/P1 ticket that is actively being worked.
- Any Status Page incident that has not been closed.
Handover Checklist
Before leaving your shift, complete the following and post it in #mb-outage:
1. Current Status
- What is happening right now? (service affected, region, clients)
- Is the issue getting better, worse, or stable?
2. Actions Already Taken
- Who has been notified? (CloudOps, Tier-2, Leadership, CSMs)
- Has the Status Page been tripped? What was posted?
- What has been tried or ruled out so far?
3. Open Items & Next Steps
- What is the incoming agent expected to do next?
- Is there a pending update due? When?
- Are there any follow-up calls or commitments already made to clients?
4. Bridge & Channel Links
- Share the active bridge link again.
- Link to the relevant #mb-outage thread.
- Link to the active Status Page incident.
5. Contact Handoff
- Confirm with the incoming agent that they have joined the bridge.
- Introduce the incoming agent on the bridge call so all teams are aware of the handover.
Handover Template:
Important: Do not leave the bridge until the incoming agent has confirmed they are on the call and understand the situation. A verbal or Slack confirmation is required before handing off.
10. Post-Outage Closure
When the outage is resolved:
- Confirm resolution with the owning service team.
- Update the Status Page to Resolved.
- Post a final closure message in #mb-outage.
- Document the following:
- Timeline of events
- Root cause summary (from Engg/CloudOps)
- Client impact summary
- Follow-up actions (if any)
If the outage spanned multiple shifts, the agent who closes the incident is responsible for collecting the full timeline from all involved agents before posting the final documentation.
11. Golden Rules
| Rule | Meaning |
|---|---|
| Never ignore PagerDuty | Every alert is actionable until dismissed. |
| Bridge starts in 3–5 minutes | Do not wait to confirm. |
| Notify CloudOps & Engg immediately | Escalation delays increase client impact. |
| Leadership PagerDuty is for awareness | Keeps leadership informed early. |
| Status Page after 15 minutes | Only after impact is confirmed. |
| Update every 30 minutes | Keeps everyone aligned. |
| Handover must be explicit | Never assume the next agent knows the context. |
| Document the timeline | Required for RCA. |
| Own the bridge until resolved | Tier-1 drives coordination — do not abandon. |
12. Ready-to-Use Templates
Slack / Internal First Response
We have detected a potential outage. Bridge started: <bridge link> CloudOps and relevant teams have been notified. Current known impact: <service / region / client>. Leadership has been notified. Next update in 30 minutes or sooner.
Internal 30-Minute Update
Outage Update — [Time] Status: [Investigating / Identified / Monitoring / Resolved] Impact: <service / region / client> Current Action: <what CloudOps or Engg is doing> Status Page: [Updated / Not yet updated] Next update at: [Time] or sooner on any change.
SOP – CO | Tier-1 Outage Handling Guidelines | Support Team
Comments
0 comments
Please sign in to leave a comment.