Technical topic
LLM judge evaluation
Design task-specific rubrics, expert reference sets, error measures, stress tests, and calibration records for model judges.
Direct answer
A language-model judge is useful only when its verdict has a defined job, an accountable reference standard, and measured failure behavior. This book follows a fictional support assistant through rubric design, corpus selection, executable scoring, miss and false-alarm analysis, adversarial stress, pipeline identity, and recalibration. It shows how expert disagreement and changing criteria constrain the claims a score can support.
What this topic helps you decide
write an llm judge rubric for one decision
Define the specific decision and its evidence boundary before writing a prompt.
measure judge misses and false alarms
Choose thresholds from the costs of misses and false alarms.
recalibrate a judge without rewriting history
Tie every verdict to a judge, rubric, corpus, and calibration revision.
Practical questions answered
- How do I write an LLM judge rubric for one decision?
Define a judge rubric with an exact decision, permitted evidence, disqualifying errors, and a human escalation state.
- What should I do when experts disagree on judge labels?
Separate ambiguous criteria, missing evidence, and legitimate expert judgment before creating a reference set for an LLM judge.
- How do I build a judge corpus that exposes serious misses?
Sample ordinary and high-consequence cases separately so an LLM judge test set reveals failures hidden by prevalence and easy examples.
- How do I measure judge misses and false alarms?
Build a confusion matrix for a defined judge decision and report serious misses, false alarms, prevalence, and uncertainty separately.
- How do I stress-test an LLM judge against shortcuts?
Test whether judge verdicts follow the rubric when answer order, length, confidence language, or irrelevant formatting changes.
- What can an LLM judge score actually support?
Attach a judge score to its rubric, corpus, uncertainty, and intended decision so it is not mistaken for universal answer quality.
- How do I recalibrate a judge without rewriting history?
Version judge, rubric, corpus, threshold, and historical verdicts so recalibration can improve current decisions without erasing earlier evidence.
Go deeper with a field guide
LLM-as-a-Judge in Practice
A Complete Guide to Expert-Led Evaluation
Explore LLM-as-a-Judge in PracticeReusable resources
The practical guides include original decision tables, schemas, or diagnostic checklists where a reusable artifact improves the answer.
