- Published on
Regression Testing Across Prompt and Model Changes
- Authors

- Name
- Mehdi Akiki
Article · Measurement
Changing an AI model is not the same as changing one library version. The model may follow instructions differently, choose tools in another order, use more context, refuse more often, or improve average quality while regressing one important customer workflow.
Prompt changes have the same problem. A sentence added to fix one failure can weaken another case.
I do not compare candidates by reading a few attractive outputs. I run a versioned regression matrix where every result points to the exact prompt, model, tools, retrieval snapshot, configuration, cases, and graders that produced it.
The system under test is larger than the prompt
An AI feature's behaviour depends on:
rendered system and user prompts
model identifier and provider version
sampling and token settings
tool names, descriptions, and schemas
authorization policy
retrieval corpus, index, filters, and top-k
conversation/history construction
application code and retry logic
grader versions and thresholds
If I record only prompt_v12, I cannot reproduce a run after the model alias or retrieval index changes.
For every candidate I create an immutable manifest:
{
"candidate": "support-agent-2026-09-22.2",
"promptHash": "sha256:...",
"model": "provider/model-version",
"parameters": { "temperature": 0, "maxTokens": 1200 },
"toolsHash": "sha256:...",
"policyVersion": "auth-17",
"corpusSnapshot": "kb-2026-09-21",
"retrievalConfig": "hybrid-rerank-6",
"evalSet": "support-regression-31",
"graderSet": "graders-9"
}
The hash covers the fully rendered instruction and examples, not only a template filename. Environment-specific additions must be visible too.
Freeze a baseline before creating the candidate
I run the currently deployed configuration and the candidate against the same case set in the same evaluation window.
This paired comparison controls many sources of noise:
baseline configuration × case 1..N × repeated trials
candidate configuration × case 1..N × repeated trials
I randomize execution order so a temporary provider or dependency slowdown does not affect only one side. When possible, I pin retrieval inputs or run against the same corpus snapshot.
The baseline is not assumed perfect. It represents behaviour users currently receive and makes regressions concrete.
Use a matrix to isolate the changed layer
If I change both prompt and model, one A/B comparison tells me whether the package improved, not which component caused it.
For important migrations I use a small factorial matrix:
| Configuration | Old prompt | New prompt |
|---|---|---|
| Old model | Baseline | Prompt-only effect |
| New model | Model-only effect | Combined candidate |
I keep tools, retrieval, and graders fixed for this comparison.
This reveals interaction. The new prompt may help the old model but hurt the new one, or the model may need a simpler instruction. I do not assume one prompt is portable across models.
When several other layers change, I do not build an enormous Cartesian product. I first test the complete candidate against baseline, then run targeted ablations for the slices that moved.
Regression cases are contracts, not exact paragraphs
Model output varies. My cases declare:
- required facts and citations;
- forbidden claims;
- allowed and forbidden tools;
- required authorization and confirmation steps;
- expected durable outcome;
- acceptable trajectory alternatives;
- latency, token, and cost budgets;
- number of trials and pass threshold.
Deterministic checks grade structure, identifiers, permissions, and effects. Rubric graders handle meaning where needed. Golden Tests for Non-Deterministic AI Outputs explains this graded assertion model.
I preserve known fixed failures as high-pass-rate regression cases. Capability cases remain difficult and measure room for improvement. Combining them into one average makes both less useful.
Repeat trials and preserve paired outcomes
One run can be lucky. I run several trials for variable cases and compare outcomes per case:
baseline: pass, pass, fail, pass, pass
candidate: pass, fail, fail, pass, fail
The candidate clearly became less stable on this case even if a different easy case raises the global average.
For a binary pass/fail suite, I inspect paired flips:
baseline fail → candidate pass: improvement
baseline pass → candidate fail: regression
both pass or both fail: unchanged outcome
With enough cases, a paired test or bootstrap confidence interval can help distinguish signal from sampling noise. I still read the individual regressions. Statistical confidence does not decide whether losing one authorization case is acceptable.
Report by risk and slice
An overall score is a summary, not a release decision. I slice by:
- task and customer workflow;
- language and input length;
- easy, normal, and adversarial cases;
- read-only and effectful tools;
- permission boundary;
- answerable and unanswerable retrieval queries;
- provider error and recovery path;
- latency and cost band;
- known production failure category.
A candidate can improve 3% overall while regressing every long-context case or every request requiring clarification.
I predeclare critical slices and thresholds before looking at results. Otherwise it is easy to explain away the failures of a candidate I already want to ship.
Keep graders fixed during the comparison
Changing the judge together with the model under test makes the score difficult to interpret. The new judge may simply prefer the new model's writing style.
I pin grader prompts, models, parsing, and thresholds for baseline-versus-candidate runs. If the grader must change, I first recalibrate it against the held-out human-labelled set and rerun both baseline and candidate with the same new grader.
For hard safety rules I use deterministic checks. A model judge cannot average away an unauthorized tool call.
Calibrating an LLM Judge Against Human Disagreement covers position bias, confidence, and human agreement.
Evaluate retrieval separately
A model migration often arrives beside embedding, reranker, or context-window changes. I preserve retrieval output for one diagnostic run:
same labelled evidence → old generator
same labelled evidence → new generator
Then I benchmark old and new retrievers independently before testing the complete RAG systems.
This prevents a stronger generator from hiding worse retrieval by answering from memorized knowledge. Evaluate Retrieval Separately From Answer Generation provides the benchmark design.
Compare trajectories, not only answers
A new model may reach the right answer using more calls or an unsafe path. I compare:
- unauthorized reads and writes;
- duplicate effects;
- tool-choice and argument errors;
- recovery after timeouts;
- calls started after cancellation;
- calls, tokens, latency, and monetary cost;
- final claims without matching effect records.
The trajectory evaluation runs beside the outcome grader. Correct final prose cannot compensate for an unsafe intermediate effect.
A two-speed suite keeps feedback practical
Full repeated evaluation can be expensive and slow. I use two levels.
Change-time sentinel suite
A small set runs on every prompt, tool-schema, or orchestration change:
- one format torture case;
- one permission boundary;
- one effect with idempotent retry;
- one unanswerable question;
- one long-context case;
- a few failures previously caused by this component.
It gives fast feedback but cannot approve a production migration alone.
Release suite
The complete representative dataset runs with repeated trials, all critical slices, cost and latency measurement, and calibrated graders. Model changes and large prompt changes require this suite.
Scheduled evaluation also catches provider drift when a model alias changes behaviour without a repository commit.
Anthropic's current agent evaluation guide recommends continuous regression suites alongside capability evaluation and running automated evals during model upgrades.
The versioned result matrix
For every run, I retain a compact record:
| Field | Why I need it |
|---|---|
| Candidate manifest hash | Reproduce the complete system configuration |
| Case and dataset version | Know which requirement was tested |
| Trial seed/order metadata | Investigate variation and execution bias |
| Per-grader evidence | Debug a failed score |
| Tool/effect trace | Detect unsafe successful runs |
| Latency, tokens, and cost | Prevent operational regressions |
| Baseline paired result | See improvements and regressions per case |
| Human override and reason | Feed calibration and future cases |
I do not retain sensitive raw inputs without an approved purpose and retention period. Production-derived cases are sanitized before entering the permanent suite.
A release gate I can defend
A candidate must satisfy rules such as:
zero new authorization or data-exposure violations
zero duplicate irreversible effects
no regression in critical workflow pass thresholds
overall quality improvement is larger than measured uncertainty
no important slice falls outside its allowed margin
p95 latency, token use, and cost remain inside budget
every material regression is reviewed and accepted explicitly
rollback configuration is tested and ready
I do not require the candidate to win every case. A model can make real trade-offs. I require those trade-offs to be visible and deliberate.
Roll out with identity and rollback
Offline evaluation reduces risk but does not reproduce every production input. I release behind a versioned configuration or feature flag:
1% canary → inspect safety and operational signals
10% → compare matched product outcomes
50% → confirm capacity and cost
100% → retain old manifest for immediate rollback
Every production trace records the candidate manifest ID. If a user reports a failure, I can reproduce the configuration and add a sanitized regression case.
Rollback includes prompt, model, tools, retrieval configuration, and policy as one tested unit. Reverting only the prompt while leaving a new tool schema can create another untested combination.
Research on regression testing for evolving LLM APIs shows why prompt effectiveness can change as models evolve. The practical response is not to freeze the system forever; it is to compare versioned systems with repeatable evidence.
My practical rule
I treat a prompt or model candidate as a complete software artifact. It has an immutable manifest, paired baseline run, versioned evaluation set, repeated trials, stable graders, risk slices, and rollout identity.
This lets me answer more than “the new model looks better.” I can say where it improved, where it regressed, how stable the result is, what it costs, whether its tool path stayed safe, and exactly which system was tested.