Incident response — page, change, notes

When checkout-api pages, the incident is in PagerDuty, the change is in GitHub, and the RCA is in a folder. Wire those two sources, drop notes you already wrote. If the digest cannot name all three, it says so. Slack waits. Create a workspace when you are ready.

Duty Clerk holds a wrist cuff to a labeled hospital bed. A phone on the nightstand is marked out.

Who it's for

SRE lead and incident commander — on-call, postmortems, deploy risk.

The problem

Incidents, fixes, and discussion threads span PagerDuty, GitHub, and Slack. Chat copilots lose links after the fire; repeat incidents rediscover the same RCA paths.

Outcome

Live incident pulses join the change pulse and the notes you already wrote. Digest cites incident + PR/deploy + RCA path, or names the miss. Private evals on your facts — no MTTR percent.

Multi-horizon BI

AI SRE-style multi-horizon reliability BI: live incident pulse, short-term change correlation, long-term RCA patterns under governance.

Real-time

PagerDuty pages and deploy events as live reliability heartbeats on dept.engineering.* streams. Slack waits (parked).

Short-term

Hours–days of ordered stream history and link-enriched incident→PR→ticket chains for blast-radius and recent recurrence.

Long-term

Optional local MIT memory stores institutional RCA narratives and similar-incident patterns agents can recall after the fire is out.

2026 AI SRE guidance stresses multi-source investigation and institutional memory—not log chat alone. Measure private evals on your incidents—not public lift claims.

Workflow map (2026)

When checkout-api pages, the incident is in PagerDuty, the change is in GitHub, and the RCA is in a folder. Wire those two sources, drop notes you already wrote. If the digest cannot name all three, it says so. Slack waits. Not a PagerDuty replacement.

SRE / ICHead of PlatformOn-call lead

Build

PagerDuty and GitHub as operational products with link enrich. Slack waits (parked).

Steer

Incident and health tools scoped to your tenancy.

Compound

Linked PRs and on-call pages become institutional memory after the fire — not one-off chat. Slack waits (parked) — no Slack Exists memory.

Learning loop

Linked PRs and on-call pages become institutional memory after the fire — not one-off chat. Slack waits (parked) — no Slack Exists memory.

Questions agents answer

  • Which PR or deploy is linked to this PagerDuty page?
  • Trace blast radius from this PagerDuty page to recent deploys.
  • What runbook and prior incidents match this failure signature?

How it works on I/O Mesh

  1. Step 1

    Ingest on-call and engineering signals

    PagerDuty and GitHub webhooks publish to dept.engineering.events.* as the Stage 0 spine. Zendesk is catalog-only — optional when live, not a ships-today peer with those two. Slack waits (parked).

  2. Step 2

    Link enrich across tables

    Cross-table advisories chain incident → PR → ticket without blocking publish — fail-open sidecar keeps hot-path ingest live.

  3. Step 3

    local MIT memory recall

    Department-scoped local MIT memory on your machine — needles and haystack evals prove agents find the right evidence.

  4. Step 4

    MCP operator tools

    trace_incident and summarize_health tools give copilots governed, department-scoped RCA paths with audit lineage.

Departments

  • Engineering
  • Ops
  • Support
Operational meshKnowledge mesh

dept.* stream patterns

  • dept.engineering.events.pagerduty
  • dept.engineering.events.github
  • dept.support.events.zendesk
Browse data product catalog →

What ships today

  • PagerDuty and GitHub connectors — Slack waits (parked). Zendesk is catalog-only / optional when live — not a Stage 0 ships-today peer. Catalog listing is not Connected.
  • Link enrich incident→PR→ticket
  • Private evals — no dual_write in buyer copy

Governed MCP tools

  • trace_incident
  • summarize_health
  • compose_operational_mesh
Typical company stage
enterprise
Evaluation path
Wire PagerDuty and GitHub, drop notes you already wrote. If the digest cannot name all three, it says so. Slack waits. Private evals on your facts — no MTTR percent. Create a workspace when you are ready.

Outcome metrics

Leading indicator
Time to linked PR evidence
Lagging indicator
Similar-incident recall and RCA completeness (private evals — no MTTR percent)

Recommended connectors

  • GitHub
  • PagerDuty

Each live connector bills as a usage meter — install from Integrations after signup.

Related solutions

Start compounding this workflow — create a workspace with the use case pre-selected.

Create a workspace Explore platform Browse all use cases