Mehdi Akiki
Published on

Regression Testing Across Prompt and Model Changes

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Article · Measurement

Changing an AI model is not the same as changing one library version. The model may follow instructions differently, choose tools in another order, use more context, refuse more often, or improve average quality while regressing one important customer workflow.

Prompt changes have the same problem. A sentence added to fix one failure can weaken another case.

I do not compare candidates by reading a few attractive outputs. I run a versioned regression matrix where every result points to the exact prompt, model, tools, retrieval snapshot, configuration, cases, and graders that produced it.

The system under test is larger than the prompt

An AI feature's behaviour depends on:

rendered system and user prompts
model identifier and provider version
sampling and token settings
tool names, descriptions, and schemas
authorization policy
retrieval corpus, index, filters, and top-k
conversation/history construction
application code and retry logic
grader versions and thresholds

If I record only prompt_v12, I cannot reproduce a run after the model alias or retrieval index changes.

For every candidate I create an immutable manifest:

{
  "candidate": "support-agent-2026-09-22.2",
  "promptHash": "sha256:...",
  "model": "provider/model-version",
  "parameters": { "temperature": 0, "maxTokens": 1200 },
  "toolsHash": "sha256:...",
  "policyVersion": "auth-17",
  "corpusSnapshot": "kb-2026-09-21",
  "retrievalConfig": "hybrid-rerank-6",
  "evalSet": "support-regression-31",
  "graderSet": "graders-9"
}

The hash covers the fully rendered instruction and examples, not only a template filename. Environment-specific additions must be visible too.

Freeze a baseline before creating the candidate

I run the currently deployed configuration and the candidate against the same case set in the same evaluation window.

This paired comparison controls many sources of noise:

baseline configuration × case 1..N × repeated trials
candidate configuration × case 1..N × repeated trials

I randomize execution order so a temporary provider or dependency slowdown does not affect only one side. When possible, I pin retrieval inputs or run against the same corpus snapshot.

The baseline is not assumed perfect. It represents behaviour users currently receive and makes regressions concrete.

Use a matrix to isolate the changed layer

If I change both prompt and model, one A/B comparison tells me whether the package improved, not which component caused it.

For important migrations I use a small factorial matrix:

ConfigurationOld promptNew prompt
Old modelBaselinePrompt-only effect
New modelModel-only effectCombined candidate

I keep tools, retrieval, and graders fixed for this comparison.

This reveals interaction. The new prompt may help the old model but hurt the new one, or the model may need a simpler instruction. I do not assume one prompt is portable across models.

When several other layers change, I do not build an enormous Cartesian product. I first test the complete candidate against baseline, then run targeted ablations for the slices that moved.

Regression cases are contracts, not exact paragraphs

Model output varies. My cases declare:

  • required facts and citations;
  • forbidden claims;
  • allowed and forbidden tools;
  • required authorization and confirmation steps;
  • expected durable outcome;
  • acceptable trajectory alternatives;
  • latency, token, and cost budgets;
  • number of trials and pass threshold.

Deterministic checks grade structure, identifiers, permissions, and effects. Rubric graders handle meaning where needed. Golden Tests for Non-Deterministic AI Outputs explains this graded assertion model.

I preserve known fixed failures as high-pass-rate regression cases. Capability cases remain difficult and measure room for improvement. Combining them into one average makes both less useful.

Repeat trials and preserve paired outcomes

One run can be lucky. I run several trials for variable cases and compare outcomes per case:

baseline: pass, pass, fail, pass, pass
candidate: pass, fail, fail, pass, fail

The candidate clearly became less stable on this case even if a different easy case raises the global average.

For a binary pass/fail suite, I inspect paired flips:

baseline fail → candidate pass: improvement
baseline pass → candidate fail: regression
both pass or both fail: unchanged outcome

With enough cases, a paired test or bootstrap confidence interval can help distinguish signal from sampling noise. I still read the individual regressions. Statistical confidence does not decide whether losing one authorization case is acceptable.

Report by risk and slice

An overall score is a summary, not a release decision. I slice by:

  • task and customer workflow;
  • language and input length;
  • easy, normal, and adversarial cases;
  • read-only and effectful tools;
  • permission boundary;
  • answerable and unanswerable retrieval queries;
  • provider error and recovery path;
  • latency and cost band;
  • known production failure category.

A candidate can improve 3% overall while regressing every long-context case or every request requiring clarification.

I predeclare critical slices and thresholds before looking at results. Otherwise it is easy to explain away the failures of a candidate I already want to ship.

Keep graders fixed during the comparison

Changing the judge together with the model under test makes the score difficult to interpret. The new judge may simply prefer the new model's writing style.

I pin grader prompts, models, parsing, and thresholds for baseline-versus-candidate runs. If the grader must change, I first recalibrate it against the held-out human-labelled set and rerun both baseline and candidate with the same new grader.

For hard safety rules I use deterministic checks. A model judge cannot average away an unauthorized tool call.

Calibrating an LLM Judge Against Human Disagreement covers position bias, confidence, and human agreement.

Evaluate retrieval separately

A model migration often arrives beside embedding, reranker, or context-window changes. I preserve retrieval output for one diagnostic run:

same labelled evidence → old generator
same labelled evidence → new generator

Then I benchmark old and new retrievers independently before testing the complete RAG systems.

This prevents a stronger generator from hiding worse retrieval by answering from memorized knowledge. Evaluate Retrieval Separately From Answer Generation provides the benchmark design.

Compare trajectories, not only answers

A new model may reach the right answer using more calls or an unsafe path. I compare:

  • unauthorized reads and writes;
  • duplicate effects;
  • tool-choice and argument errors;
  • recovery after timeouts;
  • calls started after cancellation;
  • calls, tokens, latency, and monetary cost;
  • final claims without matching effect records.

The trajectory evaluation runs beside the outcome grader. Correct final prose cannot compensate for an unsafe intermediate effect.

A two-speed suite keeps feedback practical

Full repeated evaluation can be expensive and slow. I use two levels.

Change-time sentinel suite

A small set runs on every prompt, tool-schema, or orchestration change:

  • one format torture case;
  • one permission boundary;
  • one effect with idempotent retry;
  • one unanswerable question;
  • one long-context case;
  • a few failures previously caused by this component.

It gives fast feedback but cannot approve a production migration alone.

Release suite

The complete representative dataset runs with repeated trials, all critical slices, cost and latency measurement, and calibrated graders. Model changes and large prompt changes require this suite.

Scheduled evaluation also catches provider drift when a model alias changes behaviour without a repository commit.

Anthropic's current agent evaluation guide recommends continuous regression suites alongside capability evaluation and running automated evals during model upgrades.

The versioned result matrix

For every run, I retain a compact record:

FieldWhy I need it
Candidate manifest hashReproduce the complete system configuration
Case and dataset versionKnow which requirement was tested
Trial seed/order metadataInvestigate variation and execution bias
Per-grader evidenceDebug a failed score
Tool/effect traceDetect unsafe successful runs
Latency, tokens, and costPrevent operational regressions
Baseline paired resultSee improvements and regressions per case
Human override and reasonFeed calibration and future cases

I do not retain sensitive raw inputs without an approved purpose and retention period. Production-derived cases are sanitized before entering the permanent suite.

A release gate I can defend

A candidate must satisfy rules such as:

zero new authorization or data-exposure violations
zero duplicate irreversible effects
no regression in critical workflow pass thresholds
overall quality improvement is larger than measured uncertainty
no important slice falls outside its allowed margin
p95 latency, token use, and cost remain inside budget
every material regression is reviewed and accepted explicitly
rollback configuration is tested and ready

I do not require the candidate to win every case. A model can make real trade-offs. I require those trade-offs to be visible and deliberate.

Roll out with identity and rollback

Offline evaluation reduces risk but does not reproduce every production input. I release behind a versioned configuration or feature flag:

1% canary → inspect safety and operational signals
10%      → compare matched product outcomes
50%      → confirm capacity and cost
100%     → retain old manifest for immediate rollback

Every production trace records the candidate manifest ID. If a user reports a failure, I can reproduce the configuration and add a sanitized regression case.

Rollback includes prompt, model, tools, retrieval configuration, and policy as one tested unit. Reverting only the prompt while leaving a new tool schema can create another untested combination.

Research on regression testing for evolving LLM APIs shows why prompt effectiveness can change as models evolve. The practical response is not to freeze the system forever; it is to compare versioned systems with repeatable evidence.

My practical rule

I treat a prompt or model candidate as a complete software artifact. It has an immutable manifest, paired baseline run, versioned evaluation set, repeated trials, stable graders, risk slices, and rollout identity.

This lets me answer more than “the new model looks better.” I can say where it improved, where it regressed, how stable the result is, what it costs, whether its tool path stayed safe, and exactly which system was tested.