Mehdi Akiki
Published on

Golden Tests for Non-Deterministic AI Outputs

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Article · Measurement

I like golden tests for serializers, compilers, and generated files. A known input produces a known output, and a diff shows exactly what changed.

The same idea becomes fragile when I apply it literally to an AI feature. A useful answer can be phrased in many ways. The same model may choose a different order, a different example, or another valid tool path. If the test expects one frozen string, it can fail while the product still works.

The opposite reaction is also wrong: “The output is non-deterministic, so we cannot regression-test it.”

I keep the golden case, but replace one exact expected string with graded assertions. The test fixture describes what must be true, what must never happen, and how often the system must succeed across repeated trials.

What belongs in a golden case

For an ordinary snapshot test, the output is the oracle. For an AI evaluation, I make the contract the oracle.

A useful case contains:

case ID and purpose
input and controlled context
allowed tools and permissions
required outcome facts
forbidden claims or effects
quality rubric
maximum cost or steps
number of trials
pass threshold

This lets one case accept several valid outputs without accepting everything.

For example, a support assistant answering a cancellation question may use different words, but it must identify the correct policy, avoid inventing a refund, and link to the right next action. These are stable assertions. The exact opening sentence is not.

I grade from hard facts to softer quality

I use several layers because one grader cannot answer every question well.

1. Deterministic assertions

Code should check what code can prove:

  • valid JSON or required output fields;
  • URL belongs to an allowed domain;
  • quoted identifier exists in supplied context;
  • money amount matches a source record;
  • no forbidden tool was called;
  • tool arguments respect the tenant and limits;
  • latency, token, and step budgets were not exceeded.

These assertions are fast, cheap, and explain their failures clearly.

2. Outcome assertions

For an agent, the final text may be less important than the resulting state. I check whether the ticket was created once, the correct row changed, or the requested file exists with valid content.

This catches a fluent answer that claims success after the tool failed.

3. Reference-backed assertions

Some facts can be compared with a set rather than a string. If an answer must mention two of three approved reasons, I normalize the extracted claims and compare membership. If a retrieval result must include a particular document, I inspect document IDs before evaluating prose.

4. Rubric grading

Clarity, completeness, or tone may require a human or model grader. I give the grader a narrow rubric with observable criteria. “Is this good?” is not a reproducible instruction.

A rubric item might be:

2: explains the limitation and gives a safe next step
1: mentions the limitation but gives no actionable next step
0: hides the limitation or claims unsupported certainty

The levels describe evidence, not adjectives.

A small graded assertion library

I represent every check with the same result shape:

type Grade = {
  name: string
  score: number
  maxScore: number
  required: boolean
  evidence: string
}

type TrialResult = {
  grades: Grade[]
  latencyMs: number
  toolCalls: Array<{ name: string; args: unknown }>
}

The case policy can then say:

all required assertions must pass
total rubric score must be at least 8/10
p95 latency must stay below the product budget
at least 18 of 20 trials must pass
no trial may perform a forbidden effect

I keep required safety assertions separate from averaged quality. A high clarity score must never compensate for reading another tenant's data.

Repeated trials are part of the test

A single passing run does not describe a variable system. I run important cases several times and keep the individual results.

Suppose a case passes 9 times out of 10. That can mean very different things:

  • a stable 90% behaviour;
  • a model version changed halfway through the run;
  • one retrieved document is intermittently unavailable;
  • the first request warms a cache;
  • only a particular paraphrase triggers the failure.

I record model identifier, prompt version, tool versions, retrieval snapshot, and configuration with every trial. Otherwise I know a regression happened but cannot reproduce its conditions.

For a new suite, I begin with enough trials to expose obvious variance without making every development run expensive. High-risk cases receive more trials in release checks. The right number depends on the failure rate I need to detect; it is not one magic constant.

Do not hide variance behind temperature zero

Lower temperature may reduce one source of variation, but it does not make a complete AI system deterministic. Hosted model implementations change, retrieval results move, tool responses vary, concurrency changes order, and the model can still have multiple valid next tokens.

I use stable settings for regression tests, but I test the product behaviour I actually deploy. If production allows tool retries and changing data, a perfect offline string comparison gives false confidence.

Build cases from failures without copying private data

The strongest golden cases often begin with a real failure. I do not simply paste a production conversation into the repository.

I extract the failure contract:

given two customers with similar names
when the request omits the account number
the system must ask for clarification
and must not call the payment tool

Then I create synthetic names, identifiers, amounts, and text that preserve the reasoning difficulty. This makes the case reviewable, shareable, and less likely to leak user data. Turning Production Traces Into an Evaluation Set gives the full sanitization process.

Separate capability tests from regression tests

I keep two related suites:

  • capability cases ask whether the system can solve representative and difficult tasks;
  • regression cases preserve failures we already fixed and contracts we refuse to break.

A capability score can move gradually. A regression case should normally become harder to remove than to add. If requirements change, I update the case with an explanation and review, not only because a new model fails it.

This protects the suite from becoming a scoreboard that is silently rewritten until every release looks green.

How I investigate a failed case

A combined score is useful for a dashboard but bad for debugging. I preserve the trace:

input → retrieved context → model decision → tool arguments
      → tool result → later decisions → final output → grades

Then I classify the first meaningful divergence:

  • wrong or missing context;
  • correct context ignored;
  • invalid plan;
  • unauthorized or malformed tool call;
  • tool failure handled badly;
  • correct work but misleading final answer;
  • grader disagreement.

This matters because “the case failed” does not tell me whether to change retrieval, policy, a tool contract, the prompt, or the grader.

Test the graders too

A model grader is another non-deterministic component. I keep a human-labelled calibration set containing clear passes, clear failures, and difficult disagreements. I check the grader against this set whenever its model or prompt changes.

For subjective criteria, I may also randomize answer order, hide model identity, and run the judge more than once. If a criterion can become deterministic code, I move it out of the model grader.

Anthropic's current guide to evaluating AI agents makes the same useful separation between outcomes, trajectories, graders, repeated trials, and human calibration. It also recommends starting with a small set of cases taken from real failures before scaling the suite.

A release policy I can explain

I do not reduce the suite to “average score is higher.” A release rule can be:

zero safety invariant regressions
zero duplicate side effects
no more than 1% absolute drop on core capability cases
no statistically unclear drop hidden by too few trials
latency and cost remain inside their budgets
all newly fixed production failures have regression cases

The exact thresholds belong to the product risk. What matters is declaring them before seeing the new model's result.

My practical definition of a golden test

For a non-deterministic AI output, “golden” does not mean one sacred paragraph. It means a reviewed, versioned example of acceptable behaviour.

I freeze the input, context, permissions, invariants, rubric, and evaluation conditions. I allow wording and valid reasoning paths to vary. I repeat the case enough to observe instability, and I never let an average quality score erase a hard safety failure.

This gives me regression tests that behave like engineering controls instead of screenshots of one lucky generation. The next problem is making sure a model-based grader represents human judgment, which needs its own human-anchored calibration experiment.