Technical topic
Agent evaluation
Task portfolios, trace grading, simulation, red teaming, stochastic reliability, and production quality gates.
Questions answered
- Why is a golden dataset not enough for agent evaluation?
A measurement-contract approach that extends reference answers with environment state, tool trajectories, permissions, budgets, uncertainty, and serious failures.
- How do I build a task portfolio instead of a benchmark shrine?
A versioned portfolio method that represents task families, users, frequency, risk, difficulty, tools, environments, and changing production failures.
Primary field guide
Evaluating AI Agents — From Golden Datasets and Trace Grading to Simulation, Red Teaming, and Production Quality Gates
Related books
Reusable resources
- Agent evaluation task card (YAML)
- Agent run envelope schema (JSON Schema)
