Question-led guide · how-to
How do I measure judge misses and false alarms?
Build a confusion matrix for a defined judge decision and report serious misses, false alarms, prevalence, and uncertainty separately.
Direct answer
Define the positive class from the action: for example, a positive result means “send to human review.” Compare the judge with a versioned expert reference and count true positives, false positives, false negatives, and true negatives. Report miss and false-alarm rates with denominators, case mix, and uncertainty. Overall agreement alone can be high when a rare harmful class is missed; threshold selection must reflect the cost of each error.
Define positive by the operational action
The word positive can mean a good answer or a detected problem. Ambiguity reverses precision and recall. Write “positive means send to review” at the top of the worksheet. Then freeze the rubric, reference labels, judge build, and threshold before counting. A revised criterion changes the meaning of every cell even if the model output stays the same.
Preserve all four outcomes
A false negative is a harmful case the judge permits. A false positive is an acceptable case the judge sends to review. Both have costs, but one may be much more serious. Track abstentions separately; silently treating them as pass can hide uncertainty, while treating all as fail can overwhelm reviewers. Use the same case unit across the matrix, such as one complete support answer with its evidence.
Four harmful answers slip through
Suppose a fictional evaluation contains 20 disqualifying answers and 80 acceptable ones. The judge flags 16 harmful answers and 8 acceptable answers. Its harmful-case recall is 16/20, or 80%; it misses four. Its review precision is 16/24, about 67%. A single overall agreement figure would conceal how many harmful promises reached the pass path.
Fill the decision worksheet
Keep raw counts beside every rate.
| Reference / judge | Review | Pass |
|---|---|---|
| Disqualifying | 16 true positives | 4 misses |
| Acceptable | 8 false alarms | 72 true negatives |
For this invented sample, the false-alarm rate is 8/80, or 10%. The deployment decision still needs uncertainty estimates and the actual production prevalence.
Compare thresholds at a fixed review budget
If a lower threshold catches another severe miss but doubles human review, state that trade-off explicitly. Show curves or tables by case family rather than only an optimized single number. A threshold chosen on the same examples used for final evaluation is optimistic. Recheck the chosen operating point on held-out cases and under the review capacity the team can actually provide.
Report the limits of the estimate
A small challenge set can reveal a serious failure without precisely estimating the production rate. Record confidence intervals or at least count denominators and the sampling method. If the reference labels are contested, separate label uncertainty from judge error. The worksheet supports a bounded decision only for the tested rubric, case population, and time window.
Evidence boundary for judge error rates
- scikit-learn model evaluation: scikit-learn documents classification metrics and their definitions. The library does not decide which error matters to this product.
- NIST AI evaluation toolbox: NIST discusses statistical validity of AI evaluation results. It does not provide confidence bounds for the fictional counts.
All counts and rates are constructed arithmetic; they are not measured performance for a real judge.
Evidence
Precision, recall, and confusion matrices use distinct denominators.
scikit-learn documents classification metrics and their definitions.
Primary source · official-doc · checked Oct 7, 2026
Limit: The library does not decide which error matters to this product.
Evaluation uncertainty and validity require attention to assumptions and data.
NIST discusses statistical validity of AI evaluation results.
Primary source · paper · checked Oct 7, 2026
Limit: It does not provide confidence bounds for the fictional counts.
Limitations
This worksheet needs a reliable reference, representative sampling, uncertainty analysis, and decision-specific error costs before release use.
FAQ
- Why is overall agreement insufficient?
- It can be dominated by common easy cases while the judge misses a rare high-consequence class.
- What happens to abstentions?
- Record them separately and specify whether they trigger review; do not silently fold them into pass or fail.
Related guides
Continue within LLM judge evaluation, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
