The first ten minutes of an incident,
run the same way every time
The first ten minutes of an incident are usually the same checklist every time, restart the service, pull recent logs, check the last deploy, notify the right channel, and whoever is on call ends up doing it manually under pressure at 3am. We build an agent that runs your documented runbook the moment an alert fires, so a human arrives to diagnostics already gathered instead of a blank terminal.
The process today
Most teams with any operational maturity have written runbooks: documented steps for the incidents that happen often enough to be predictable, a service that occasionally needs a restart, a queue that backs up under a specific condition, a database connection pool that exhausts under a known pattern. The problem is not that the runbook does not exist, it is that executing it well under pressure, at an inconvenient hour, while still half awake, is inconsistent even for an experienced on-call engineer, and far more so for someone newer to the rotation who has read the runbook once in onboarding and is now seeing the real incident for the first time.
The second cost is the time spent just gathering context before any actual fix starts: pulling the recent logs, checking what deployed in the last few hours, confirming which service is actually affected versus which one is just downstream noise. None of that is hard, but doing it by hand takes minutes that matter when customers are affected, and it is the same handful of steps almost every time.
The third is inconsistency in the incident record itself. What actually happened, in what order, gets reconstructed from memory afterward for the postmortem, which means the account is usually incomplete or slightly wrong by the time it is written down.
What the agent does
Your documented runbooks get converted into steps the agent can actually execute: which logs to pull, which metrics to check, which deploys to look at, and which specific, pre-approved actions are safe to run automatically, restarting a known-flaky service, clearing a stuck queue, scaling a resource that is under load. The moment a matching alert fires, the agent gathers that diagnostic context immediately, runs whatever safe first-response steps your team has explicitly approved, creates the incident channel, and populates it with the gathered context before a human even joins.
For anything beyond the pre-approved safe actions, the agent stops at a clear handoff point and pages a person with everything gathered so far: what the symptoms are, what changed recently, what the runbook says to try next, and what it has already ruled out. If the incident does not match any documented pattern, it says so plainly rather than guessing, pages a human immediately, and flags the gap for your team to turn into a new runbook afterward. Every action taken is logged in order with a timestamp, building an accurate timeline ready for the postmortem instead of one reconstructed from memory. Typical integrations: PagerDuty or Opsgenie for alerting, Slack or Telegram for the incident channel, and your existing logging and monitoring stack for diagnostics.
What stays with humans
Deciding the actual fix for anything beyond the narrow, pre-approved safe actions stays with the engineer who takes the handoff; the agent’s job is to make sure they start that work with full context instead of losing the first ten minutes to gathering it. Which actions get automatic execution rights is a decision your team makes explicitly, runbook by runbook, before launch, never something the agent decides for itself mid-incident. The postmortem itself, what to actually change afterward, is a team discussion the agent’s timeline supports but does not replace.
Guards
Every action the agent takes during an incident is logged in sequence with a timestamp, building a precise record for the postmortem. Automatic execution is scoped to a specific, pre-approved list of safe, reversible actions; anything outside that list always goes to a human, with no path for the agent to improvise beyond its approved scope. A kill switch disables automatic action execution in one message while keeping diagnostic gathering and paging running, useful if a specific automated step is suspected of causing problems.
Price and timeline
| Option | Price | What it covers | Timeline |
|---|---|---|---|
| Single automation | from $1,200 | A small set of your most common incident types, diagnostic gathering, safe first-response actions | 1 to 3 weeks |
| Department package | from $2,800 | Incident runbook execution plus on-call summary reports and exception triage | 3 to 5 weeks |
Running cost is usually $20 to $60 a month in model usage depending on incident frequency.
Related
This pairs well with on-call summary reports to close the loop after the incident is resolved, and with exception triage and assignment for the issues that led up to it. For the public-facing side of a live incident, see status page and incident updates. Full package details are on the AI agents service page and the automation-everything overview; for how we run incident response on infrastructure we operate ourselves, see the secure infrastructure case study and the factory ERP recovery case study.
Tired of every incident starting with the same ten minutes of manual digging? Get in touch and we will turn your runbooks into something an agent can run.
Tired of doing this by hand? We can take the whole routine off your team, not just this step: Routine takeover, from $400 →
FAQ
Does the agent actually fix incidents on its own?
For a narrow set of safe, reversible, pre-approved actions, like restarting a known-flaky service or clearing a stuck queue, yes. Anything beyond that, it gathers diagnostics, runs the safe first steps, and hands off to a person with full context, it does not improvise a fix.
How much does incident runbook automation cost?
From $1,200 for a small set of your most common incident types, live in 1 to 3 weeks. A fuller runbook library across multiple services usually runs $2,000 to $3,000.
What if the incident does not match any documented runbook?
The agent says so explicitly, gathers whatever general diagnostics it can, pages a human immediately, and flags the gap so your team can document a runbook for that case afterward.
Which alerting and incident tools does this connect to?
PagerDuty, Opsgenie, or a Slack or Telegram-based on-call setup, connected to your logging and monitoring stack, whatever that already is.
Who decides what counts as a safe automatic action?
Your team, explicitly, runbook step by runbook step, before anything is automated. Nothing gets automatic execution rights without your engineers reviewing and approving that specific step first.