Cover of AI Cost Engineering by Leo J. Li

Signal Studio field guide

AI Cost Engineering

Token Economics, Model Routing, Caching, Capacity, and Cost per Outcome

An engineering guide to measuring and controlling the full cost of useful AI outcomes across requests, retries, tools, reviews, routing, caching, and capacity.

For: AI platform engineers, FinOps practitioners, engineering leaders

Status
Live
Format
Kindle eBook
ASIN
B0HG4CW692
Page updated

What this book helps you do

This book replaces cost-per-token dashboards with an outcome ledger that follows every attempt through acceptance, rejection, reversal, and human review. It connects token demand, model routing, caching, capacity purchases, serving efficiency, and evaluation gates so teams can reduce waste without quietly moving cost into failures or degraded product quality.

Problems this book helps you solve

  • Token spend is visible, but nobody can state the cost of one accepted outcome.
  • Retries, tool calls, human review, and reversals disappear from unit economics.
  • A cheaper model looks successful because quality and escalation costs are measured elsewhere.
  • Caching is enabled without proving reuse, privacy boundaries, or invalidation behavior.
  • Provisioned capacity is purchased from average demand while tail latency and burst shape are ignored.
  • Teams optimize provider invoices but cannot reconcile them with product and business metrics.

Decisions you will be able to make

  • How to define an outcome unit, acceptance window, and denominator that resists gaming.
  • Which request, tool, review, platform, and reversal costs belong to an outcome.
  • When a smaller model should answer, escalate, refuse, or route to a larger model.
  • What to cache, where to scope it, and when reuse is unsafe or economically irrelevant.
  • When on-demand, batch, priority, provisioned, or self-hosted capacity fits the workload.
  • Which cost and quality thresholds should block a release.

Who this book is for

  • Teams with meaningful AI spend that cannot connect invoices to accepted product outcomes.
  • Platform groups comparing model, routing, caching, and capacity policies under the same quality contract.
  • Engineering and finance partners building a shared AI unit-economics language.

Who this book is not for

  • Readers seeking a static provider price comparison that will remain current without verification.
  • Teams prepared to reduce nominal cost without measuring quality, serious failures, or user outcomes.

Reading path

  1. Define the outcome unitName the accepted result, attribution window, failure states, and reversal policy before computing a rate.
  2. Build the request-to-outcome ledgerRetain every model call, tool call, retry, review, and shared-cost allocation.
  3. Control token demandTreat context, output, reasoning, and repeated work as separate engineering levers.
  4. Route with evidenceCompare policies by task family and quality gate, including escalation and refusal behavior.
  5. Cache the right objectModel reuse, scope, invalidation, privacy, hit quality, and counterfactual savings.
  6. Choose and operate capacityMatch purchasing and serving strategies to workload shape, service levels, and failure capacity.

Cost is a property of a completed workflow

A model request is one component of an AI product. A useful cost model follows the work through retries, retrieval, tools, guardrails, review, failure, and reversal. That makes cost comparable to an outcome that product and finance teams can recognize.

Optimize policies, not isolated prices

The book treats routing, caching, capacity, and evaluation as policies with measurable consequences. A lower unit price is valuable only when the resulting policy preserves the required outcome distribution and serious-failure boundary.

Use it with

Start with the public cost-per-outcome ledger, then replace the synthetic row values with an approved local outcome definition and reconciled cost sources.

Evidence and method

Provider prices and product capabilities are treated as dated inputs that require re-verification, not permanent constants. Worked numbers are labeled synthetic. The reusable framework is the measurement contract: define the outcome, include all attempts, hold quality constant, expose allocation choices, and retain uncertainty instead of manufacturing precision.

Read a sample

Signal Studio does not reproduce manuscript chapters on this site. Open the Amazon listing to use Read Sample or Kindle Instant Preview

Resources

Errata and related guidance

Report or review an erratum.

English editorial review: Codex native-English editorial review, .