Technical topic

LLM judge evaluation

Design task-specific rubrics, expert reference sets, error measures, stress tests, and calibration records for model judges.

Direct answer

A language-model judge is useful only when its verdict has a defined job, an accountable reference standard, and measured failure behavior. This book follows a fictional support assistant through rubric design, corpus selection, executable scoring, miss and false-alarm analysis, adversarial stress, pipeline identity, and recalibration. It shows how expert disagreement and changing criteria constrain the claims a score can support.

What this topic helps you decide

write an llm judge rubric for one decision

Define the specific decision and its evidence boundary before writing a prompt.

measure judge misses and false alarms

Choose thresholds from the costs of misses and false alarms.

recalibrate a judge without rewriting history

Tie every verdict to a judge, rubric, corpus, and calibration revision.

Practical questions answered

Go deeper with a field guide

Reusable resources

The practical guides include original decision tables, schemas, or diagnostic checklists where a reusable artifact improves the answer.