Question-led guide · evaluation
How do I build a judge corpus that exposes serious misses?
Sample ordinary and high-consequence cases separately so an LLM judge test set reveals failures hidden by prevalence and easy examples.
Direct answer
Build the corpus around the decision’s error classes rather than around conveniently available prompts. Include frequent ordinary work, rare disqualifying cases, borderline evidence, missing documentation, and subgroup shifts. Record collection time, source version, labeling process, and why each case belongs. Keep a held-out portion untouched during prompt tuning, then report serious misses with their own denominators instead of burying them in aggregate agreement.
Sample for the decision, not the easiest collection path
A production log may contain many straightforward approved answers and few documented severe errors. Random sampling alone can therefore produce a high-looking score while revealing little about the failure the judge is meant to prevent. Define the unacceptable mistake and seek cases that test it directly. Keep prevalence estimates separate from a deliberately enriched challenge set.
Make case identity and evidence reproducible
Each case needs the input, answer, permitted evidence snapshot, rubric version, human reference process, and collection reason. Remove customer secrets or use governed synthetic substitutes where possible. A case copied from a later postmortem can leak the answer into the judge input. Label the evidence actually available at decision time and record transformations that could change interpretation.
A high agreement score hides one expensive miss
Consider a fictional set of 100 support answers: 99 routine replies and one unsupported promise of premium access. A judge that passes all 100 obtains 99% agreement if the routine labels are all pass. It fails the only disqualifying case. The team expands the challenge set with varied false promises and records the severe-miss rate separately, while still keeping a representative sample for prevalence.
Use a coverage matrix before splitting data
A small matrix makes missing families visible before test numbers appear.
| Case family | Collection route | Critical check |
|---|---|---|
| Ordinary correct | Representative production sample | False alarm |
| Known false claim | Expert-authored or confirmed case | Severe miss |
| Ambiguous evidence | Dated reference conflict | Abstention |
| Changed policy | Old and new effective dates | Version handling |
| New wording | Held-out paraphrase | Shortcut resistance |
Protect the holdout from prompt iteration
Choose and freeze a subset before tuning. If a prompt author repeatedly inspects failures and rewrites the judge, those cases have become development data. Keep a separate untouched evaluation set or refresh it under a recorded process. Report the exact corpus version and sampling weights; a weighted production estimate and an enriched challenge-set rate answer different questions.
Audit what the corpus still cannot show
A finite set cannot cover every future product rule or adversarial wording. Inspect subgroups, source types, response lengths, and missing-data conditions. When a new policy launches, add a dated case family and re-evaluate instead of assuming the old corpus transfers. A good corpus is an operating asset whose gaps remain visible, not a one-time benchmark.
Evidence boundary for judge corpora
- NIST AI evaluation toolbox: NIST analyzes evaluation validity and statistical assumptions for AI measurement. It does not prescribe this support corpus composition.
- Judging LLM-as-a-Judge research: The paper examines agreement and bias in model judges on specific benchmarks. Its benchmark tasks do not determine support-case prevalence.
The 100-case arithmetic is invented and demonstrates denominator choice, not observed judge performance.
Evidence
Benchmark validity depends on the population and measurement assumptions.
NIST analyzes evaluation validity and statistical assumptions for AI measurement.
Primary source · paper · checked Oct 7, 2026
Limit: It does not prescribe this support corpus composition.
Model-judge behavior can vary with evaluation setup and bias.
The paper examines agreement and bias in model judges on specific benchmarks.
Primary source · paper · checked Oct 7, 2026
Limit: Its benchmark tasks do not determine support-case prevalence.
Limitations
Coverage is task-specific. Sampling and reference quality must be documented before a result can support a deployment claim.
FAQ
- Should rare failures be oversampled?
- Yes for a challenge set, while reporting its result separately from a representative production estimate.
- Can I reuse test cases while tuning the judge?
- Cases repeatedly inspected during tuning become development data; retain a separate untouched evaluation set.
Related guides
Continue within LLM judge evaluation, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
