Operations · Reliability

One incident.
Evidence to verified recovery.

Incident & Reliability joins Datadog monitoring, Bugsnag application evidence, incident command, governed Engineering fixes and recovery verification without collapsing their safety boundaries.

1

Unified workflow

Detect → Evidence → Respond → Fix → Verify

1

Detect and assess

A live reliability monitor correlates provider signals to incidents on a recurring loop (about every 60s); you can also sync active Datadog monitors or record a bounded manual signal. Cards show severity, service, environment, release, fleet distribution, spike count and affected users when the provider reports them.

2

Correlate evidence

Datadog and Bugsnag signals can join one open incident when service, environment, release, team and timing cross a conservative evidence threshold. Weak matches remain separate; accepted matches record their score and reasons.

3

Coordinate response

Assign an owner and incident commander from the configured Jira project roster. Advance the governed lifecycle with evidence and follow the incident-response playbooks retrieved from the dedicated RAG collection.

4

Route the work

Code-shaped evidence creates a read-only Engineering investigation. Host, Agent and configuration evidence records an Operations handoff. Neither action patches, deploys or resolves automatically.

5

Verify recovery

A patch or PR is not resolution. Confirm Datadog health, Bugsnag error reduction and the configured stable recovery window before resolving, then complete the review and corrective actions.

2

FABRIC Intelligence

Grounded hypotheses, never automatic truth

Bounded provider evidence

Forge retains normalized facts and authoritative links—not raw Datadog logs or complete Bugsnag payloads. Provider credentials remain server-side and provider data stays authoritative in its source.

Incident-management RAG

Analysis retrieves diversified chunks only from the incident_response collection. The result separates observations from likely causes, missing evidence, verification steps, response recommendations and cited playbooks.

Impact integrity

Fleet groups, request spikes and affected users are different units. Missing spike or user values remain “Not reported”; FABRIC must not convert alerting hosts into users or invent impact.

Human authority

FABRIC analysis is advisory. Lifecycle transitions, ownership, mitigation, Engineering plans, patches, approval, publication and resolution remain explicit governed decisions.

3

Governed fixes

Incident context follows the Engineering run

A forge.incident-handoff.v1 package carries up to eight bounded correlated signals, severity, service, environment, release and source references into Observe and Ground. Datadog and Bugsnag evidence enters as the internal trust tier (vetted systems), distinct from untrusted free-form text. The Engineering state machine remains independent — the handoff feeds the full 17-state governed engine (observed → grounded → planned → patching → patched → reviewing → reviewed → validating → validated → awaiting_approval → approved → committed → draft_pr_created, then release observation → verification), which stops at a draft PR.

Do not resolve on patch creation.Resolution requires deployed recovery evidence and a stable monitoring window. A blocked investigation is a valid governed outcome, not permission to bypass missing evidence or validation.
4

Setup

Configure once, route per workflow

Datadog

Save one shared service-access token under Settings → Connectors. Dev, Product and Production reuse it; each workflow keeps its own enable flag, monitor tags, environments, Jira project and recovery window.

Knowledge

Register incident communication, escalation, response, checklist, LiveOps and on-call pages in the incident_response collection, ingest them, and confirm source freshness before relying on citations.

5

Shared language

Severity model and response expectations

SEV 1 through SEV 4 use one governed definition across teams. The incident detail panel includes a collapsible legend plus FABRIC recommendations for severity, impact, ownership and next action. Read the full Incident Severity Model →