Use case: incident response runbooks
Runbooks go stale, and an AI agent let loose on a production incident without one is a risk nobody wants to take. With Floxar, the runbook becomes a flow that your engineers and your agents both run: the agent does the fast, repeatable diagnosis, and the on-call engineer takes over before anything risky happens.
The parts involvedโ
| Layer | In this case |
|---|---|
| Trigger | An alert from your monitoring system |
| Agent | Your incident agent, which starts the run and works the early steps |
| Process (Floxar) | The incident flow: diagnose, match a known cause, apply or escalate, resolve, close |
| Systems | Monitoring, logs and metrics, the service catalogue, the deployment tool |
| Credentials | Read access for diagnostics, kept in the Secrets Vault |
| People | The on-call engineer, through the on-call queue |
How a run goesโ
- The alert starts a trail. Your agent, or an integration you run, receives the alert and starts the incident flow, recording the alert's details on the first step.
- The agent collects diagnostics. Each diagnostic step names the check to run and what to record: error rates, recent deployments, the health of dependencies. The agent runs them with read-only credentials and records the results.
- The agent matches a known cause. A decision step lists the documented causes and the evidence for each. When the evidence matches one, the agent takes that path and its documented fix.
- Risky actions go to a person. Restarts, rollbacks and anything else that cannot be safely undone are steps for the on-call engineer. The agent writes a hand-off note with what it found and transfers the trail to the on-call queue.
- The engineer resolves and closes. The engineer sees every diagnostic result on the trail, carries out the fix, and closes the run with a closing reason such as "resolved by rollback".
When nothing matches a known cause, the flow sends the run to the engineer straight away, with the diagnostics already gathered.
What you get from the recordโ
Every incident leaves a trail: which checks ran, what they returned, which cause was matched, who acted and when. Across many incidents, the trails and analytics show which branches of the runbook fire most, where time goes, and which incidents always end with a person. Improve the flow, and the next incident runs the improved version, for the agent and the engineers alike.
For the general pattern, see Designing AI workloads with Floxar.
Last reviewed: 2026-10-06