P1 vs P2 Incidents: A Field Guide to Classification That Engineers Actually Follow

Most P1 vs P2 frameworks get ignored because they're written for auditors, not on-calls. Here's a classification model your team will actually use during a 3 AM page.

TL;DR — The Two-Question Framework

  1. Is a customer-facing surface degraded right now (or will it be in 15 min)?
    → Yes = P1 or P2
  2. Can the on-call fix this alone in under an hour?
    → No = P1 · Yes = P2 · No customer impact = P3

Every incident management vendor ships a severity matrix. Almost none of them survive contact with a real on-call. The classic 5x5 grid — impact on the X axis, urgency on the Y axis — looks tidy on a slide and falls apart at 2 AM when the on-call has 90 seconds to decide whether to wake up the VP of Engineering. This guide replaces the matrix with a decision tree your team will actually follow, anchored to two questions: who is feeling pain right now, and how fast is it spreading.

Why most severity frameworks fail

The first failure mode is too many tiers. A five-level model (P1 through P5) sounds rigorous; in practice, only P1 and P2 are ever paged on, P3 becomes a dumping ground, and P4–P5 are functionally a backlog. The second failure mode is overloaded criteria — a definition that requires the on-call to evaluate revenue impact, customer count, regulatory exposure, brand risk, and time-of-day before assigning a tier. The third failure mode is no enforcement loop: the team picks a severity, ships the incident, and nobody ever reviews whether it was right.

The result is severity drift. Engineers default to P2 for almost everything because it's safe — high enough to get attention, low enough to not embarrass anyone. Real P1s get classified down because the on-call doesn't want to wake the executive. Minor P3s get classified up because the engineer wants visibility. Within a year the data is meaningless and the SLA dashboard tells leadership a story disconnected from reality.

The two-question framework

Cut the matrix. Ask two questions in order. Question one: is a customer-facing surface degraded right now, or will it be within 15 minutes if we do nothing? If yes, you're at P1 or P2. Question two: can the team I'd page right now actually fix this in under an hour with the people on call, or does this need executive escalation, vendor engagement, or cross-team coordination? If it needs escalation, P1. If the on-call can fix it, P2.

Everything else — internal tools, batch failures, monitoring gaps that haven't yet caused customer pain — is P3 or below and goes through the normal ticket flow, not the pager. That's the entire framework. Two questions, two tiers worth paging on, one bucket for everything else.

The Classification Decision Tree

QuestionAnswerSeverity
Customer-facing surface degraded?NoP3 or below
Customer-facing surface degraded?Yes → needs exec/cross-teamP1
Customer-facing surface degraded?Yes → on-call can fix aloneP2

🔴 P1 — Customer Pain + Escalation Needed

Definition: Company-wide problem the on-call alone cannot solve in 1 hour.

Examples:

  • ❌ Checkout down for all customers
  • ❌ Primary database failing with data-loss risk
  • ❌ Security incident in progress
  • ❌ Region-wide cloud outage

SLA Targets:

ActionTarget
Responder acknowledgment5 minutes
Customer acknowledgment15 minutes
Status page updateEvery 30 minutes
Mitigation target4 hours
Public RCA5 business days

Default Response:

  • War-room bridge call within 5 minutes
  • Executive on the line within 15 minutes
  • Customer status page updates every 30 minutes

🟡 P2 — Customer Pain On-Call Can Handle

Definition: Single service degraded, on-call has access and runbook to fix.

Examples:

  • ⚠️ One microservice returning 5xx for 5% of requests
  • ⚠️ Search latency degraded but not down
  • ⚠️ Non-critical admin tool fully down
  • ⚠️ Message queue processing delayed 15 min

SLA Targets:

ActionTarget
Responder acknowledgment30 minutes
Internal status update1 hour
Mitigation target8 hours
RCA10 business days

Default Response:

  • Page the on-call
  • Post status update internally
  • No executive escalation unless 1hr passes

🟢 P3 and Below — Disciplined Backlog

Definition: No current customer-facing degradation. Ticket, owner, target date.

Examples:

  • ✅ Monitoring noise without customer impact
  • ✅ Non-customer-facing degradation
  • ✅ Issues found during business hours
  • ✅ Single pod OOMing with auto-restart masking

SLA Targets:

ActionTarget
Acknowledgment4 business hours
Fix target5 business days
Public RCANot required

Real Incident Examples

ScenarioSeverityWhy
Payment provider 30% errors, checkout drops to 70%, customer complaints floodingP1Customer-facing, needs exec + vendor
Single K8s pod OOMing hourly, auto-restart masking, no customer impactP3No customer pain
One of three message queue consumers stuck, 15-min delay, on-call has restart fixP2Customer-facing but on-call can fix

❌ Common Mistakes to Avoid

  • Don't classify on raw percentages — 5% errors on 100 RPS ≠ 5% on 10K RPS
  • Don't classify based on time of day — severity is severity; response can flex, tier shouldn't
  • Don't let responder grade own homework — every P1/P2 gets next-day classification review

How to Enforce Classification Consistency

3 controls that catch 90% of drift:

  1. Post-incident pop-up — "Was this classified correctly?" at incident close. Unsures get reviewed.
  2. Weekly cohort review — 5 minutes per incident, 3 people, every P1/P2 from prior week.
  3. Quarterly recalibration — 30 random incidents, re-classify blind, measure agreement. Below 80% = framework drift.

SLA Summary by Severity

TierResponder AckCustomer AckMitigationRCA
P15 min15 min4 hours5 business days
P230 min1 hour8 hours10 business days
P34 business hoursN/A5 business daysNot required

Bind these to your platform's SLA engine so the timers run automatically and the dashboard tells the truth.

Conclusion

Severity is a decision-making tool, not a documentation exercise. Two questions, two paging tiers, one disciplined backlog. Run the framework for one quarter, measure agreement, and you'll see paging volume drop, RCA quality rise, and on-call burnout fall. The matrix can stay on the wiki for the auditors; the team will use the tree at 2 AM.

If you want to automate the classification workflow, enforce SLAs, and track the prevention work that keeps P1s from recurring, try SLAShield's major incident workflow. Start with a free trial or see pricing for your team size.

Frequently Asked Questions

Q: What is the difference between P1 and P2?

P1 requires executive escalation and cross-team coordination — the on-call alone cannot resolve it in under an hour. P2 is customer-facing but the on-call has the access and runbook to fix it independently.

Q: When should you escalate a P2 to P1?

Escalate when: the on-call has been working for 30+ minutes without progress, customer impact is spreading, a second service is degrading, or leadership asks for an update.

Q: How do you prevent severity drift?

Three controls: post-incident classification review, weekly P1/P2 cohort review, and quarterly blind recalibration. Below 80% agreement signals framework drift.

Q: Should P3 incidents have RCAs?

No formal RCA required, but every P3 needs an owner and a close date. Unowned P3 tickets are closed at the weekly review. High-volume P3 patterns may warrant a lightweight review.