Question-led guide · evaluation
How do I evaluate AIOps before putting it on call?
Build an AIOps evaluation portfolio that tests decision quality, severe misses, operator workload, data gaps, and rollback before on-call adoption.
Direct answer
Evaluate the decision the on-call team will make, not a model metric alone. Use a portfolio of ordinary incidents, multi-fault cases, quiet periods, data loss, deployment changes, and high-consequence failures. Compare the proposed workflow with the existing process for detection, actionability, severe misses, review effort, and recovery. Introduce it as advice under human control, then widen pager dependence only with a measured stop and rollback rule.
State exactly what on-call would delegate
An AIOps tool can prioritize an alert, suggest a group, draft a diagnosis, or trigger an action. These are different promises. Specify the proposed operator behavior and the point at which a human can decline or correct it. A system that only drafts a hypothesis may need less evidence than one whose silence suppresses a page. The evaluation must mirror the real consequence.
Use cases that expose the expensive mistake
Historical incidents are useful but often underrepresent concurrent faults and missing telemetry. Add controlled cases for a silent detector, an over-merged incident, a forecast during a traffic shift, and a recovery with an unknown effect. Retain case provenance and ensure the system does not see the answer through postmortem notes during evaluation. Grade serious failure modes separately from an average score.
A clean replay misses the collector failure
Imagine a test set of 120 past incidents in which the AIOps assistant highlights the correct service for 108. All cases have complete telemetry. In a constructed live drill, a collector queue fills, trace coverage falls, and the assistant confidently declares no impact. The 90% historical result did not measure the observation gap that matters to the pager. The next portfolio version includes this failure.
Record the actual decision impact
Use a scorecard that compares the current workflow with the proposed one.
| Measure | What to inspect |
|---|---|
| Detection | Significant events and time to actionable page |
| Severe miss | Impact hidden or wrong recovery proposed |
| Operator cost | Review, correction, and duplicate pages |
| Evidence | Source and missing-data visibility |
| Recovery | Stop control and return to previous workflow |
Release advice before authority
Run the assistant beside the existing on-call path first. Operators see the proposal and its evidence, while the original alerting remains authoritative. Review disputed outputs and actual corrections; keep a record of model, rule, data, and policy revisions. Only after the decision-level evidence meets the agreed threshold should any page suppression or automatic action receive separate approval and a limited rollout.
Keep a withdrawal rule visible
Define what failure stops the pilot: an unobserved serious incident, a repeated bad grouping, loss of input coverage, or an unreviewable recommendation. Rehearse rollback before broadening exposure and verify that the old workflow still functions. Retain the exact evaluation and deployment identities so an incident can be traced to the version that influenced the operator.
Evidence boundary for AIOps pager readiness
- Google SRE alerting on SLOs: Google SRE gives decision-relevant dimensions for alert evaluation. Those dimensions do not by themselves certify an AIOps assistant.
- NIST AI Risk Management Framework: NIST frames governance, mapping, measurement, and management as linked risk functions. The framework does not supply thresholds for this fictional pager portfolio.
The incident counts are illustrative. A real readiness decision requires representative local cases and a documented risk threshold.
Evidence
Alerting quality includes precision, recall, detection time, and reset time.
Google SRE gives decision-relevant dimensions for alert evaluation.
Primary source · official-doc · checked Oct 7, 2026
Limit: Those dimensions do not by themselves certify an AIOps assistant.
AI risk governance and measurement must continue through use.
NIST frames governance, mapping, measurement, and management as linked risk functions.
Primary source · official-doc · checked Oct 7, 2026
Limit: The framework does not supply thresholds for this fictional pager portfolio.
Limitations
This portfolio is a starting worksheet. Case weights, incident labels, severe-failure thresholds, and rollback mechanics must be set by the operating team.
FAQ
- Can strong historical accuracy justify suppressing pages?
- No. Page suppression needs evidence about significant misses, telemetry gaps, and live withdrawal behavior at the actual consequence level.
- Should a pilot begin with automated recovery?
- A safer first pilot presents inspectable advice under the existing on-call process, then evaluates separate authority for effects.
Related guides
Continue within AIOps alerting and operations, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
