Question-led guide · how-to

How do I write an LLM judge rubric for one decision?

Define a judge rubric with an exact decision, permitted evidence, disqualifying errors, and a human escalation state.

Direct answer

Start with the decision that will consume the verdict, such as whether a support answer contains an unsupported product claim. Define the evidence the judge may use, observable pass and fail criteria, and an abstain state for missing context. Give experts contrasting examples and adjudicate disagreements before automating. Version the rubric with its owner; a fluent model explanation cannot repair an ambiguous standard.

A rubric boundary connects one product decision to permitted evidence, pass criteria, disqualifying errors, and abstention.
Rubric scope: The boundary is an author-designed rubric model; it does not establish that a particular judge follows the criteria. This is an author-created explanatory model, not measured system evidence.

Name the action the verdict will influence

A judge cannot be calibrated for “good writing” in the abstract when its score will block a release. State whether it triages review, recommends a revision, or makes a pass/fail gate. A support team may care most about a false product promise; style and warmth can be recorded separately. The higher-consequence decision needs the clearest error definition and an accountable owner.

Make evidence eligibility part of the rule

List what the judge may inspect: customer request, approved product documentation, proposed answer, and perhaps a dated policy note. Exclude hidden release notes or later customer outcomes if the production judge cannot see them. Describe how to handle absent, conflicting, or outdated references. If the evidence set is unbounded, experts can reach different verdicts without either being careless.

The polished answer with one false promise

In a fictional support case, an answer accurately explains five setup steps but promises that every account includes a feature available only in a premium tier. A broad helpfulness score might pass it. The rubric says an unsupported entitlement claim is a disqualifying error even when the rest is useful. If the tier documentation is missing, the correct state is “review needed,” not an invented pass.

Write a compact decision sheet

The sheet turns a preference into an inspectable rule.

Field Specific entry
Decision Send for human review or permit normal release
Evidence Versioned entitlement documentation and answer
Fail condition Unsupported feature availability claim
Abstain Evidence absent or conflicting
Owner Product policy expert and evaluation lead

Test disagreement before prompting a model

Ask two experts to label the same mixed cases independently. Compare their rationales and identify whether the wording, evidence, or underlying policy caused disagreement. Rewrite the criterion and preserve the earlier rubric version. The goal is not to pressure every reviewer into agreement; it is to make the boundary and legitimate exceptions visible enough that a machine verdict can be evaluated.

Keep the output narrower than the rubric

Require a verdict, cited evidence identifier, triggered criterion, and uncertainty state. Do not let a model produce a confident reason unsupported by the allowed materials. Review both false passes and false alarms against expert decisions. A clear rubric is necessary for evaluation, but the actual judge still needs task-specific calibration on held-out cases.

Evidence boundary for judge rubrics

  • NIST AI Risk Management Framework: NIST frames AI risks through governance, mapping, measurement, and management. It does not supply the support-answer rubric in this example.
  • Judging LLM-as-a-Judge research: The paper studies model-judge agreement and several judgment biases. Its benchmark results cannot validate this support policy.

The entitlement case is invented. Domain experts must approve the real policy and labeled examples before deployment.

Evidence

  1. AI use needs explicitly governed purpose and measured risks.

    NIST frames AI risks through governance, mapping, measurement, and management.

    Primary source · official-doc · checked Oct 7, 2026

    Limit: It does not supply the support-answer rubric in this example.

  2. Model judges show task-dependent biases and require evaluation against human preferences.

    The paper studies model-judge agreement and several judgment biases.

    Primary source · paper · checked Oct 7, 2026

    Limit: Its benchmark results cannot validate this support policy.

Limitations

The worksheet defines one decision; it does not prove model reliability, policy correctness, or transfer to another support domain.

FAQ

Can the model write its own acceptance policy?
It can suggest wording, but accountable experts must define and approve the rule that will govern a consequential decision.
Why include an abstain verdict?
Missing or conflicting permitted evidence makes a binary pass or fail misleading; abstention routes that case to review.

Continue within LLM judge evaluation, or use one of these adjacent diagnostics:

Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.