Question-led guide · governance
How do I recalibrate a judge without rewriting history?
Version judge, rubric, corpus, threshold, and historical verdicts so recalibration can improve current decisions without erasing earlier evidence.
Direct answer
Create a new calibration revision rather than editing old verdicts in place. Keep the judge build, rubric, reference set, threshold, input evidence, and decision time for each historical result. Evaluate the candidate on matched old and new cases, including serious failures, then shadow it before promotion. If policy changes, label the new standard explicitly and preserve what earlier decisions meant under the old one.
A verdict belongs to a complete revision
Store the model identifier, prompt, tool or retrieval configuration, rubric, threshold, source evidence, and output schema with every judgment. A version name without immutable contents is weak: two runs can use the same label while loading different policy text. Historical decisions should be reproducible enough to explain what the team knew and which rule was applied, even if the underlying model is no longer available.
Separate model drift from policy change
A model may become less reliable under new user language while the acceptance standard stays fixed. A policy owner may also change the standard itself. These call for different comparisons. For drift, evaluate old and candidate judges against one stable reference. For policy change, create a new reference view and explain why historical labels cannot be compared as if they measured the same target.
A corrected rubric changes a dashboard line
In a constructed release history, version R3 allowed a narrow exception for beta feature language. Version R4 removes it. A dashboard that rescored last month’s cases under R4 shows a jump in “judge errors,” but no past judge behavior changed. The team keeps the R3 decision record, adds R4 counterfactual scores in a separate series, and dates the policy transition.
Keep a calibration ledger
A ledger joins each result to the conditions that produced it.
| Identity | Record |
|---|---|
| Judge package | Model and prompt digests |
| Standard | Rubric and policy effective date |
| Evidence | Corpus and source revisions |
| Decision rule | Threshold and abstain handling |
| Outcome | Verdict, review, and later correction |
| Promotion | Owner, date, and rollback target |
Compare candidates on matched cases
Run old and new judge revisions on the same frozen evidence. Include routine work, severe misses, false alarms, contested cases, and policy-transition examples. Inspect where a candidate flips a verdict and why. Do not choose solely on aggregate improvement if a high-consequence family regresses. Reserve a holdout for the final decision and record the reviewer who accepted the change.
Promote with a shadow and a reversal plan
A shadow phase records candidate verdicts without changing the production action. Compare them with human decisions and observe review workload. Once promoted, retain a fast rollback to the previous complete revision and watch for unfamiliar case types. A rollback restores the decision system; it does not erase actions already taken under the promoted judge, which require separate review if they were consequential.
Evidence boundary for judge recalibration
- NIST AI Risk Management Framework: NIST frames measurement and management as continuing functions. It does not define this judge version ledger.
- Model Cards for Model Reporting: The model-card paper establishes a reporting pattern for model context and performance. A model card alone cannot reproduce historical verdicts or policy state.
The R3/R4 history is illustrative. Real recalibration requires immutable artifacts and retained decision records.
Evidence
AI risk management includes ongoing measurement and management.
NIST frames measurement and management as continuing functions.
Primary source · official-doc · checked Oct 7, 2026
Limit: It does not define this judge version ledger.
A model report should capture intended use, evaluation, and limitations.
The model-card paper establishes a reporting pattern for model context and performance.
Primary source · paper · checked Oct 7, 2026
Limit: A model card alone cannot reproduce historical verdicts or policy state.
Limitations
The ledger is a recording design. Reproduction may still be limited by retired model access, stochastic behavior, and unavailable external evidence.
FAQ
- Should I rescore historical cases after a rubric change?
- Yes for a labeled counterfactual comparison, while preserving the original verdict and the standard that governed it.
- Does rolling back the judge undo its past decisions?
- No. It restores a prior rule for future decisions; consequential past actions need their own review and correction.
Related guides
Continue within LLM judge evaluation, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
