Operations · Reliability
One incident.
Evidence to verified recovery.
Incident & Reliability joins Datadog monitoring, Bugsnag application evidence, incident command, governed Engineering fixes and recovery verification without collapsing their safety boundaries.
1Unified workflow
Detect → Evidence → Respond → Fix → Verify
1Detect and assess
A live reliability monitor correlates provider signals to incidents on a recurring loop (about every 60s); you can also sync active Datadog monitors or record a bounded manual signal. Cards show severity, service, environment, release, fleet distribution, spike count and affected users when the provider reports them.
2Correlate evidence
Datadog and Bugsnag signals can join one open incident when service, environment, release, team and timing cross a conservative evidence threshold. Weak matches remain separate; accepted matches record their score and reasons.
3Coordinate response
Assign an owner and incident commander from the configured Jira project roster. Advance the governed lifecycle with evidence and follow the incident-response playbooks retrieved from the dedicated RAG collection.
4Route the work
Code-shaped evidence creates a read-only Engineering investigation. Host, Agent and configuration evidence records an Operations handoff. Neither action patches, deploys or resolves automatically.
5Verify recovery
A patch or PR is not resolution. Confirm Datadog health, Bugsnag error reduction and the configured stable recovery window before resolving, then complete the review and corrective actions.
2FABRIC Intelligence
Grounded hypotheses, never automatic truth
Bounded provider evidence
Forge retains normalized facts and authoritative links—not raw Datadog logs or complete Bugsnag payloads. Provider credentials remain server-side and provider data stays authoritative in its source.
Incident-management RAG
Analysis retrieves diversified chunks only from the incident_response collection. The result separates observations from likely causes, missing evidence, verification steps, response recommendations and cited playbooks.
Impact integrity
Fleet groups, request spikes and affected users are different units. Missing spike or user values remain “Not reported”; FABRIC must not convert alerting hosts into users or invent impact.
Human authority
FABRIC analysis is advisory. Lifecycle transitions, ownership, mitigation, Engineering plans, patches, approval, publication and resolution remain explicit governed decisions.
3Governed fixes
Incident context follows the Engineering run
A forge.incident-handoff.v1 package carries up to eight bounded correlated signals, severity, service, environment, release and source references into Observe and Ground. Datadog and Bugsnag evidence enters as the internal trust tier (vetted systems), distinct from untrusted free-form text. The Engineering state machine remains independent — the handoff feeds the full 17-state governed engine (observed → grounded → planned → patching → patched → reviewing → reviewed → validating → validated → awaiting_approval → approved → committed → draft_pr_created, then release observation → verification), which stops at a draft PR.
Do not resolve on patch creation.Resolution requires deployed recovery evidence and a stable monitoring window. A blocked investigation is a valid governed outcome, not permission to bypass missing evidence or validation.
4Setup
Configure once, route per workflow
Datadog
Save one shared service-access token under Settings → Connectors. Dev, Product and Production reuse it; each workflow keeps its own enable flag, monitor tags, environments, Jira project and recovery window.
Knowledge
Register incident communication, escalation, response, checklist, LiveOps and on-call pages in the incident_response collection, ingest them, and confirm source freshness before relying on citations.
5Shared language
Severity model and response expectations
SEV 1 through SEV 4 use one governed definition across teams. The incident detail panel includes a collapsible legend plus FABRIC recommendations for severity, impact, ownership and next action. Read the full Incident Severity Model →