Question-led guide · governance
What evidence should an automated recovery require?
Define an automation gate that binds a recovery proposal to current state, narrow authority, expected effect, and a verifiable receipt.
Direct answer
An automated recovery needs a specific approved proposal, a fresh state check, bounded target set, valid authority, an execution identity, and an independent effect observation. Reject the action if any material part changes while it waits. Preserve a stop path and a human handoff for unknown outcomes. A successful command response is only attempt evidence; the workflow should reconcile the external system before declaring recovery or retrying.
Bind approval to the action that will run
A generic instruction to repair the service cannot authorize arbitrary later commands. Store the target resource, operation, parameters, expected state, scope, approver, expiry, and proposal digest. Immediately before execution, compare the current request with the approved one. If the target set or risk changed, return to review rather than stretching the old approval to cover a different operation.
Separate command acceptance from restored service
A tool may return success after queuing a restart while users still see errors. Another tool may time out after the change was committed. Preserve attempt IDs and inspect the target system for the actual effect. The recovery claim should include the service-level observation that mattered to users, not only a deployment or API response. Unknown outcome is a legitimate state that demands reconciliation.
The retry would restart a healthy replica
In a constructed service, an operator approves restarting one unhealthy worker at 15:00. By 15:04 the worker has recovered and the scheduler moved its workload. The queued automation still holds the earlier approval. A fresh state check rejects the restart. Later a network timeout obscures another attempt; the workflow queries the operation receipt before deciding whether a retry could duplicate the effect.
Use a compact gate card
The card is a review artifact and must be backed by real policy enforcement.
| Gate | Required record | Refusal condition |
|---|---|---|
| Proposal | Target, action, parameters, digest | Material mismatch |
| Authority | Principal, scope, expiry | Expired or broader action |
| State | Current version and health | Already recovered or changed |
| Attempt | Idempotency key and response | Unknown outcome without reconciliation |
| Effect | Target receipt and user signal | Effect unverified |
Rehearse the uncertain-response path
Test a timeout after the downstream system has applied the action. The automation must not assume failure and issue an unbounded retry. It should query the operation identity, observe target state, and either close with an effect receipt or hand the unknown case to a person. Check the actual idempotency behavior of the receiving system; a local key alone cannot enforce it.
End automation at a named boundary
Set maximum attempts, elapsed time, affected resources, and observation gaps before release. A recovery that improves one metric while expanding errors elsewhere should stop and escalate. Review all denied actions as useful evidence about the gate, not as pressure to weaken it. Keep a tested manual control available when the automation’s own dependencies are unavailable.
Evidence boundary for automated recovery
- NIST AI Risk Management Framework: NIST organizes AI risk work around govern, map, measure, and manage functions. It does not specify the execution protocol for this fictional restart.
- Google SRE canarying releases: Google SRE describes partial, time-limited release exposure and evaluation. Canary principles do not prove that a particular recovery command was safe.
The restart is hypothetical. Implement the gate in the actual authority and effect systems, then test failure and duplicate-attempt behavior.
Evidence
AI risk management needs governance, measurement, and ongoing management rather than one-time model acceptance.
NIST organizes AI risk work around govern, map, measure, and manage functions.
Primary source · official-doc · checked Oct 7, 2026
Limit: It does not specify the execution protocol for this fictional restart.
A limited rollout observes effects before widening exposure.
Google SRE describes partial, time-limited release exposure and evaluation.
Primary source · official-doc · checked Oct 7, 2026
Limit: Canary principles do not prove that a particular recovery command was safe.
Limitations
This is a design gate, not a certified automation policy. Local identity, state, idempotency, and effect verification must be implemented and rehearsed.
FAQ
- Does an idempotency key make retries safe by itself?
- No. The receiving system must implement the expected retention and duplicate semantics, and an unknown effect still needs reconciliation.
- Can a success response close the incident?
- Only after the target effect and the relevant user-facing signal have been observed at the required scope.
Related guides
Continue within AIOps alerting and operations, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
