Cover of AI Agents for SRE by Leo J. Li

Signal Studio field guide

AI Agents for SRE

Building LLM Root Cause Analysis on OpenTelemetry Traces, Logs, and Metrics

A field guide to building SRE agents that investigate incidents through explicit evidence, falsifiable hypotheses, bounded authority, and replayable evaluations.

For: site reliability engineers, observability engineers, AI platform teams

Status
Live
Format
Kindle eBook
ASIN
B0HFW35GFZ
Page updated

What this book helps you do

This book is for teams that want an agent to investigate incidents without turning telemetry into confident fiction. It shows how to define root-cause claims, join signals through defensible identities, expose observation gaps, give agents narrow investigative tools, and require evidence for every transition from symptom to hypothesis to conclusion.

Problems this book helps you solve

  • The RCA bot repeats the alert instead of explaining the failure mechanism.
  • The last error or slowest span is treated as the root cause.
  • Logs, traces, and metrics are joined by loose time windows or unstable names.
  • Missing telemetry is silently interpreted as evidence that nothing happened.
  • The agent receives a telemetry dump but cannot run targeted falsification queries.
  • Incident conclusions cannot be replayed, graded, or challenged by an operator.

Decisions you will be able to make

  • What evidence is sufficient for a root-cause claim at each risk level.
  • Which resource, trace, exemplar, and change identifiers are safe join keys.
  • Which investigative tools the agent may use and which actions remain human-owned.
  • When the agent must report unknown, request more telemetry, or hand over.
  • How to separate localization, causal mechanism, contributing factors, and impact.
  • How to construct replay packets and release gates for an RCA system.

Who this book is for

  • SRE teams designing or reviewing an AI-assisted incident investigation workflow.
  • Observability teams standardizing evidence across traces, logs, metrics, topology, and changes.
  • Platform engineers who need an auditable boundary between analysis and operational action.

Who this book is not for

  • Teams looking for a universal prompt that turns arbitrary telemetry into proven causality.
  • Organizations that have not established basic service identity, retention, or incident ownership.

Reading path

  1. Define the claimSeparate alerts, symptoms, causal candidates, contributing conditions, and defensible root-cause statements.
  2. Build the evidence planeUse OpenTelemetry resources and signal-specific semantics without pretending that collection equals explanation.
  3. Join without inventing relationshipsChoose stable identities, bounded time windows, exemplars, and explicit uncertainty for cross-signal pivots.
  4. Run hypothesis-led investigationsGive the agent tools to query, compare, falsify, stop, and hand over instead of summarizing a dump.
  5. Constrain authorityKeep investigation, recommendation, approval, and production effects as distinct control points.
  6. Replay and operateGrade evidence use, causal reasoning, abstention, latency, cost, and serious failure behavior.

Choose this book when the investigation process is the product

The central design problem is not whether a language model can produce a plausible incident narrative. It is whether the system can preserve identity, provenance, time, uncertainty, and operator authority while moving from an alert to a testable explanation.

What changes after reading

You should be able to review an RCA agent as an evidence system: identify which observations it can access, inspect every join it makes, challenge each causal step, reproduce a prior run, and decide where automation must stop.

Use it with

Pair the book with a representative incident casebook, an OpenTelemetry schema inventory, and a release suite containing both well-observed incidents and deliberate observation gaps.

Evidence and method

The book separates standards and vendor documentation from incident observations, composite teaching cases, and author-created operating frameworks. A trace establishes recorded execution structure, not causality by itself. Every case states what the available evidence supports, what remains unknown, and which additional test could change the conclusion.

Read a sample

Signal Studio does not reproduce manuscript chapters on this site. Open the Amazon listing to use Read Sample or Kindle Instant Preview

Resources

Errata and related guidance

Report or review an erratum.

English editorial review: Codex native-English editorial review, .