Question-led guide · governance

What should I do when experts disagree on judge labels?

Separate ambiguous criteria, missing evidence, and legitimate expert judgment before creating a reference set for an LLM judge.

Direct answer

Keep independent expert labels and rationales before adjudication. Classify disagreements as a rule ambiguity, evidence gap, policy conflict, or genuinely discretionary judgment. Resolve what the product decision requires, version the revised criterion, and retain contested cases. A majority vote can hide the very boundary the judge must learn; disputed examples may need abstention or specialist review instead of a forced label.

A decision tree routes independent expert labels through evidence gaps, rule ambiguity, or policy conflict into a documented record.
Reference quality: The tree is an author-created adjudication aid; agreement after discussion alone does not measure reference validity. This is an author-created explanatory model, not measured system evidence.

Preserve the first independent judgments

If reviewers talk before recording their labels, their later agreement is hard to interpret. Keep each initial verdict, evidence used, criterion cited, and a short rationale. Blind review can reduce anchoring by a senior reviewer, but it does not remove a badly specified policy. Treat the disagreement as diagnostic data about the evaluation system rather than a nuisance to average away.

Locate the source of the difference

A reviewer may lack a document that another saw, interpret the same exception differently, or apply a newer policy. These require different repairs. An evidence gap needs a corrected packet; rule ambiguity needs rewritten criteria; a legitimate discretionary choice may need a range or escalation category. Forcing all three into one binary label makes the judge target unstable.

A policy revision changes the apparent truth

In a constructed support review, two experts fail an answer promising a data-export feature; a third passes it using a later product note. The note became effective after the customer interaction. Once the team aligns the evidence snapshot to the answer date, the third judgment no longer describes the same task. The review log retains both verdicts and the reason for the corrected reference.

Use an adjudication record that survives revisions

The record should identify the exact case and rule version.

Item Preserve
Initial labels Reviewer IDs, time, evidence versions
Disagreement class Missing fact, ambiguous rule, policy conflict, discretion
Resolution Revised criterion or specialist decision
Scope Which historical cases require relabeling
Residual uncertainty Why a case remains contested

Measure reference quality, not just model fit

After clarification, relabel a sample without revealing the previous consensus. Report remaining disagreement by criterion and case family. A judge that exactly reproduces an unstable reference set may still be unsuitable for the product decision. Keep difficult cases in a separate challenge set so future rubric changes can be evaluated against the boundary they were meant to repair.

Do not erase the old standard

Attach labels to a rubric and policy revision. When the rule changes, retain the earlier labels for historical release analysis and create a new reference view for current decisions. Document whether a past verdict was wrong under its own standard or merely differs under the new one. This distinction matters when comparing model versions across time.

Evidence boundary for expert reference labels

  • NIST AI evaluation toolbox: NIST discusses statistical assumptions and validity of AI evaluation conclusions. It does not adjudicate this fictional policy dispute.
  • Model Cards for Model Reporting: The paper proposes reporting intended uses and evaluation details with models. A reporting format cannot settle an ambiguous domain policy.

The support-policy timeline is illustrative. A real program needs accountable policy owners and reviewer calibration.

Evidence

  1. Evaluation validity depends on assumptions about the measurement target and data.

    NIST discusses statistical assumptions and validity of AI evaluation conclusions.

    Primary source · paper · checked Oct 7, 2026

    Limit: It does not adjudicate this fictional policy dispute.

  2. Evaluation reporting should record intended use and performance context.

    The paper proposes reporting intended uses and evaluation details with models.

    Primary source · paper · checked Oct 7, 2026

    Limit: A reporting format cannot settle an ambiguous domain policy.

Limitations

Adjudication can improve reference clarity but does not make inherently discretionary cases objectively binary.

FAQ

Should I use majority vote for every disagreement?
No. First determine whether the disagreement comes from different evidence, criteria, or timing; a vote may conceal a broken standard.
Do I relabel old cases after a policy change?
Create a versioned current view while preserving the historical labels and the policy under which they were assigned.

Continue within LLM judge evaluation, or use one of these adjacent diagnostics:

Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.