- Published on
Offline Evals Predict; Online Signals Correct
- Authors

- Name
- Mehdi Akiki
Article · Measurement
An offline evaluation tells me how an AI system behaves on cases I selected. Production tells me which cases reality selected.
I need both.
Offline evals are repeatable, inspectable and safe to run before release. They can test dangerous failures without exposing users. But they simplify the traffic, tools, latency, history and incentives of the live product.
Online signals include real outcomes and new failure shapes. But they are noisy, delayed and biased by who used the feature and what the system chose to show. They cannot safely replace a release gate.
My working rule is simple: offline evals predict whether a candidate is ready; online signals correct what the offline suite misunderstood.
The two questions are different
Before release, I ask:
Does candidate B satisfy our known quality, safety, latency and cost contracts
better than the current system on a controlled case set?
After release, I ask:
Is the deployed system helping its real users, and what important behaviour
did our controlled case set fail to represent?
The first question needs stable inputs and paired comparisons. The second needs product instrumentation, incident evidence, user research and careful interpretation.
If I merge them into one score, I lose what each measurement can prove.
A measurement architecture I can debug
versioned candidate
│
├── offline suite ──→ release gate
│ │
│ └── per-case evidence
│
└── canary deployment ──→ production signals
│ │
│ ├── safety/incident signals
│ ├── task and product outcomes
│ ├── latency, errors and cost
│ └── sampled human review
│
└── sanitized failures ──→ candidate regression cases
Every event carries the deployed configuration ID: model, rendered prompt hash, tools, retrieval snapshot or version, policy and application release. Without that join key, I can observe a failure but cannot reliably connect it to the system that produced it.
The versioned regression matrix explains how I create that candidate identity before release.
What offline evaluation does well
An offline suite is the right place for deterministic pressure:
- known regressions that must stay fixed;
- authorization and data-exposure boundaries;
- unanswerable questions and refusal behaviour;
- injected provider timeouts and malformed tool results;
- long-context and multilingual slices;
- repeated trials for variable outcomes;
- paired cost, latency and quality comparison;
- adversarial cases that should never be sent through real user traffic.
I can inspect every failure and rerun the exact case. A candidate does not receive production traffic until its critical contracts pass.
The suite is still a sample. A 98% offline score does not mean 98% of users will succeed. It means 98% of this dataset's scored trials passed under this harness.
What production can reveal
Live measurement sees the complete system around the model:
- input distributions that were missing from the dataset;
- actual tool and provider reliability;
- user corrections, retries and abandonment;
- downstream effects rather than attractive text;
- cache, retrieval and permission behaviour at real scale;
- latency tails and capacity limits;
- workflows users invent rather than workflows I expected.
I watch operational signals immediately during a canary. Product outcomes may take longer. A generated support answer can return in two seconds, while “the issue stayed resolved” may need days of observation.
Current agent-evaluation guidance from Anthropic recommends combining evals with production monitoring, A/B tests and user research. This combination matters because an agent's path through tools and environment is part of its behaviour.
A click is not correctness
Online proxies are easy to count and easy to misunderstand.
more clicks could mean more useful results or more confusion
longer sessions could mean engagement or failure to finish
fewer escalations could mean resolution or a hidden escalation button
more acceptance could mean quality or automation bias
I connect a metric to a product hypothesis and add a counter-metric. For example:
| Hypothesis | Primary signal | Counter-signal |
|---|---|---|
| Answer resolves the task | no repeat contact in 72 hours | unsafe or incorrect review rate |
| Coding suggestion helps | accepted and survives CI | reverted within seven days |
| Agent completes workflow | verified durable outcome | duplicate or unauthorized effects |
| Retrieval gives evidence | cited claim supported | relevant source omitted |
Google's Rules of Machine Learning makes a useful distinction between the objective optimized and the broader metrics that describe system health. I apply the same caution to LLM products: one convenient proxy is not the product truth.
Online comparison needs an exposure record
If candidate B is shown to harder users than candidate A, raw success rates are not comparable. I record assignment and exposure:
{
"requestId": "req-812",
"candidate": "assistant-2026-09-25.3",
"experiment": "retrieval-canary-4",
"variant": "candidate",
"eligible": true,
"exposed": true,
"taskSlice": "account-recovery",
"outcomePending": true
}
The record distinguishes eligibility, assignment, actual exposure and later outcome. It avoids counting users assigned to a variant who never received its output. For high-risk changes I can shadow the candidate first, scoring it without showing or executing its actions.
Production feedback is selected by the system
A model changes what users see and therefore what future feedback exists. If a retrieval system never returns a document, users cannot click it. If an agent blocks a workflow early, later-step failures disappear from the logs.
This selection bias is why I retain randomized holdouts where ethical and practical, compare predeclared slices, and use targeted human review. I do not continuously train from positive clicks as if they were clean labels.
Google's production ML guidance warns about training-serving skew and feedback loops. An AI system adds more sources of skew because model, prompt, tools and retrieval can all change the observed distribution.
Turn evidence into cases safely
When production finds a meaningful failure, I do not copy the raw conversation into a permanent repository.
I first create a minimal case that preserves the failure mechanism:
production evidence
→ classify risk and root cause
→ remove or replace personal and confidential data
→ minimize the reproducer
→ obtain required review/consent
→ add expected contract and provenance
→ place in held-out or regression set
The production-trace workflow gives the privacy and retention details.
Not every poor rating becomes a regression. It may describe a product preference, missing capability, provider outage or unclear UI. I preserve the cause, not merely the unhappy output.
Use different alarms for different speeds
I split measurement by response time:
seconds: unauthorized calls, errors, latency, cost spikes
hours: sampled quality, tool completion, fallback use
days: repeated contact, reversions, retained task success
weeks: user trust, adoption, workflow and business outcomes
Fast safety signals can stop a canary automatically. Slow product metrics should not be forced into an instant launch decision. I keep the previous complete configuration ready for rollback while delayed signals mature.
My practical rule
I never ask offline or online evaluation to prove what it cannot observe.
Offline cases create a repeatable release contract. A canary limits exposure. Production signals reveal distribution, operational and product reality. Version IDs connect the evidence. Human review interprets ambiguous cases. Sanitized failures improve the next offline suite.
This is a loop, but not a blind self-training loop. Each transfer between production evidence and permanent evaluation data is an engineering and privacy decision. That is how online signals correct the prediction without turning noisy user behaviour into automatic truth.