Mehdi Akiki
Published on

Calibrating an LLM Judge Against Human Disagreement

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Article · Measurement

An LLM judge can review more outputs than a human team. This makes it useful, but it does not make its scores true.

I treat an LLM judge like a measurement instrument. Before I use it to approve a model or prompt change, I compare it against reviewed human decisions, inspect where people disagree with each other, and decide which errors are acceptable for the product.

The surprising part is that humans are not a perfect single oracle either. Two careful reviewers can interpret “complete,” “safe,” or “clear” differently. Calibration is therefore not only “make the model copy one person.” It is an experiment that reveals whether the rubric has a stable meaning.

Start with the decision the score will control

I do not begin by asking a judge to give answers from 1 to 10. I begin with the downstream decision:

  • block a release when safety regresses;
  • rank two candidate answers;
  • send uncertain outputs to human review;
  • monitor a quality trend;
  • label a failure category for investigation.

Each decision has a different tolerance for error. A noisy judge may still help discover trends. The same judge may be unacceptable as the only gate for a high-impact action.

For a first calibration experiment, pairwise comparison is often easier than an absolute score. The reviewer sees answer A and answer B for the same case and chooses:

A is better
B is better
tie / meaningfully equivalent
cannot judge from supplied evidence

This still needs a rubric, but it reduces disagreement about what a “7” means.

Build a calibration set that contains disagreement

If every example has one excellent answer and one obviously broken answer, almost any judge looks strong. I intentionally include:

  • clear wins and clear failures;
  • two good answers with different trade-offs;
  • correct answers with poor explanations;
  • fluent answers with one unsupported claim;
  • safe refusals and unnecessary refusals;
  • cases where supplied context is incomplete;
  • long and short answers;
  • outputs from different model families;
  • examples known to cause human disagreement.

I keep the set separate from the cases used to tune the judge prompt. Otherwise I can overfit the rubric and demonstrations to the final test.

Production failures are useful starting points, but I remove private data and preserve only the failure structure. The process in Turning Production Traces Into an Evaluation Set applies here too.

Human labels need a protocol

I ask at least two reviewers to label each important case independently before discussing it. They receive:

  • the user request;
  • only the context available to the evaluated system;
  • both candidate outputs in randomized order;
  • a short rubric with observable criteria;
  • an option to mark the case unclear.

I hide model names and system versions when possible. Reviewers can otherwise reward the system they expect to be stronger.

After independent labeling, I preserve both original labels. Adjudication can produce a reference decision, but replacing disagreement with one final label loses valuable information.

If humans disagree often on one criterion, I first improve the rubric or split the concept. The judge cannot reliably measure a target that the team has not defined.

The agreement experiment

Assume I have 120 answer pairs. Two humans label all pairs. A third reviewer adjudicates disagreements without seeing the judge result.

I then run the LLM judge with:

  1. answer order randomized;
  2. model identity removed;
  3. a fixed judge prompt and model version;
  4. structured output containing label, confidence, rubric evidence, and cannot_judge;
  5. a second run with A and B swapped;
  6. repeated runs on the ambiguous subset.

I calculate at least:

raw agreement = matching labels / all comparable cases
human-human agreement
judge-human agreement against each reviewer
judge-adjudicated agreement
agreement by rubric slice
swap consistency

Raw agreement is easy to explain but can be misleading when one label dominates. I also calculate Cohen's kappa for two raters, or another suitable chance-corrected measure, and show the confusion matrix. I do not hide the label distribution behind one number.

An illustrative result might look like this:

SliceCasesHuman-humanJudge-adjudicatedSwap-consistent
Clear factual correctness3597%94%97%
Unsupported claims3090%83%90%
Completeness trade-offs3077%73%80%
Refusal quality2584%68%76%

These numbers are an example of the report format, not results from a real product. The point is that one overall 82% would hide weak refusal judgments.

Position bias deserves its own test

A pairwise judge may prefer the first or second response independently from quality. Swapping answer order gives a direct signal.

For each pair:

run 1: A then B
run 2: B then A

After mapping the labels back to the original answers, both runs should agree. When they do not, I can:

  • classify the case as uncertain;
  • use repeated randomized judgments;
  • improve rubric evidence;
  • avoid letting this case control a release automatically.

Research has documented position bias in LLM judges; the paper Judging the Judges examines this problem and mitigation methods. I still measure it on my own task because behaviour depends on the judge, prompt, and answer distribution.

Evidence is more useful than a bare label

I ask the judge to connect every criterion to observable evidence:

criterion: factual support
label: B
evidence_from_context: document 3 states the 30-day limit
evidence_from_A: A claims 60 days without support
evidence_from_B: B uses the documented 30 days
confidence: high

I validate references when possible. A judge explanation can itself invent evidence, so fluent reasoning is not proof.

For deterministic checks—exact amounts, known citations, valid tool names—I use code instead of the model judge. The judge should spend its uncertainty budget on questions that genuinely require judgment.

Confidence must be calibrated, not decorative

A 0.93 emitted by a model is not automatically a 93% probability of correctness. I group judgments into confidence bands and compare them with observed agreement on held-out human labels.

For example:

Claimed confidenceCasesObserved agreement
High6092%
Medium3574%
Low2552%

Again, these are illustrative. A useful pattern would allow me to accept high-confidence low-risk decisions automatically and route the rest to review. If every label says “high,” confidence provides no routing value.

Google Research describes related work on calibrating autoraters to distributions of human preferences in Judging with Confidence.

Disagreement is a result, not noise to delete

When the judge disagrees with humans, I classify the cause:

  • judge ignored a rubric requirement;
  • judge relied on outside knowledge not present in context;
  • answer order changed its preference;
  • human reviewer missed a factual issue;
  • rubric allows two reasonable interpretations;
  • reference answer is wrong or outdated;
  • case does not contain enough evidence.

Some cases should be repaired. Some should be removed from automatic gating. Some reveal a product decision the team has not made.

I keep an unclear or cannot_judge outcome. Forcing every ambiguous case into pass or fail creates apparent precision and poor decisions.

Prevent judge overfitting

It is easy to change the judge prompt until it agrees with the calibration labels. That can improve the current spreadsheet while reducing generalization.

I split examples into:

  • development cases used to improve the rubric and prompt;
  • a held-out calibration set used to choose the judge version;
  • a later audit sample from fresh production-like outputs.

When prompt, judge model, or rubric changes, I version it and rerun the complete experiment. I compare both agreement and decision impact. A two-point agreement improvement is not useful if the judge becomes much more likely to miss the failure category that blocks releases.

The release gate I use

The acceptable threshold comes from consequences, not an internet benchmark. A possible policy is:

hard safety assertions: deterministic or human-reviewed, never judge-only
high-confidence quality cases: automatic judge gate after calibration
ambiguous or low-confidence cases: human review
judge drift sample: reviewed every release or every week
material judge version change: full recalibration

I also keep a manual override visible. If engineers override the judge repeatedly for the same reason, that is calibration data, not an annoyance to ignore.

My practical rule

I use an LLM judge to scale a defined evaluation process, not to invent one.

First I define the decision and rubric. Then humans label a difficult, representative set independently. I measure human-human agreement, judge-human agreement, label-specific errors, order sensitivity, and confidence. Finally, I choose which judgments can be automated and which still require people.

The original G-Eval paper showed that model-based evaluators can correlate better with human judgments than older automatic metrics on some text-generation tasks, while also warning about possible bias toward LLM-generated text; see G-Eval. Anthropic's guide to agent evaluations also recommends combining graders and periodically calibrating automated evaluation against humans.

The judge becomes valuable when I know its error shape. Before that, it is only another model producing a confident-looking answer.