Mehdi Akiki
Published on

Offline Evals Predict; Online Signals Correct

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Article · Measurement

An offline evaluation tells me how an AI system behaves on cases I selected. Production tells me which cases reality selected.

I need both.

Offline evals are repeatable, inspectable and safe to run before release. They can test dangerous failures without exposing users. But they simplify the traffic, tools, latency, history and incentives of the live product.

Online signals include real outcomes and new failure shapes. But they are noisy, delayed and biased by who used the feature and what the system chose to show. They cannot safely replace a release gate.

My working rule is simple: offline evals predict whether a candidate is ready; online signals correct what the offline suite misunderstood.

The two questions are different

Before release, I ask:

Does candidate B satisfy our known quality, safety, latency and cost contracts
better than the current system on a controlled case set?

After release, I ask:

Is the deployed system helping its real users, and what important behaviour
did our controlled case set fail to represent?

The first question needs stable inputs and paired comparisons. The second needs product instrumentation, incident evidence, user research and careful interpretation.

If I merge them into one score, I lose what each measurement can prove.

A measurement architecture I can debug

versioned candidate
      │
      ├── offline suite ──→ release gate
      │       │
      │       └── per-case evidence
      │
      └── canary deployment ──→ production signals
                 │                    │
                 │                    ├── safety/incident signals
                 │                    ├── task and product outcomes
                 │                    ├── latency, errors and cost
                 │                    └── sampled human review
                 │
                 └── sanitized failures ──→ candidate regression cases

Every event carries the deployed configuration ID: model, rendered prompt hash, tools, retrieval snapshot or version, policy and application release. Without that join key, I can observe a failure but cannot reliably connect it to the system that produced it.

The versioned regression matrix explains how I create that candidate identity before release.

What offline evaluation does well

An offline suite is the right place for deterministic pressure:

  • known regressions that must stay fixed;
  • authorization and data-exposure boundaries;
  • unanswerable questions and refusal behaviour;
  • injected provider timeouts and malformed tool results;
  • long-context and multilingual slices;
  • repeated trials for variable outcomes;
  • paired cost, latency and quality comparison;
  • adversarial cases that should never be sent through real user traffic.

I can inspect every failure and rerun the exact case. A candidate does not receive production traffic until its critical contracts pass.

The suite is still a sample. A 98% offline score does not mean 98% of users will succeed. It means 98% of this dataset's scored trials passed under this harness.

What production can reveal

Live measurement sees the complete system around the model:

  • input distributions that were missing from the dataset;
  • actual tool and provider reliability;
  • user corrections, retries and abandonment;
  • downstream effects rather than attractive text;
  • cache, retrieval and permission behaviour at real scale;
  • latency tails and capacity limits;
  • workflows users invent rather than workflows I expected.

I watch operational signals immediately during a canary. Product outcomes may take longer. A generated support answer can return in two seconds, while “the issue stayed resolved” may need days of observation.

Current agent-evaluation guidance from Anthropic recommends combining evals with production monitoring, A/B tests and user research. This combination matters because an agent's path through tools and environment is part of its behaviour.

A click is not correctness

Online proxies are easy to count and easy to misunderstand.

more clicks        could mean more useful results or more confusion
longer sessions    could mean engagement or failure to finish
fewer escalations  could mean resolution or a hidden escalation button
more acceptance    could mean quality or automation bias

I connect a metric to a product hypothesis and add a counter-metric. For example:

HypothesisPrimary signalCounter-signal
Answer resolves the taskno repeat contact in 72 hoursunsafe or incorrect review rate
Coding suggestion helpsaccepted and survives CIreverted within seven days
Agent completes workflowverified durable outcomeduplicate or unauthorized effects
Retrieval gives evidencecited claim supportedrelevant source omitted

Google's Rules of Machine Learning makes a useful distinction between the objective optimized and the broader metrics that describe system health. I apply the same caution to LLM products: one convenient proxy is not the product truth.

Online comparison needs an exposure record

If candidate B is shown to harder users than candidate A, raw success rates are not comparable. I record assignment and exposure:

{
  "requestId": "req-812",
  "candidate": "assistant-2026-09-25.3",
  "experiment": "retrieval-canary-4",
  "variant": "candidate",
  "eligible": true,
  "exposed": true,
  "taskSlice": "account-recovery",
  "outcomePending": true
}

The record distinguishes eligibility, assignment, actual exposure and later outcome. It avoids counting users assigned to a variant who never received its output. For high-risk changes I can shadow the candidate first, scoring it without showing or executing its actions.

Production feedback is selected by the system

A model changes what users see and therefore what future feedback exists. If a retrieval system never returns a document, users cannot click it. If an agent blocks a workflow early, later-step failures disappear from the logs.

This selection bias is why I retain randomized holdouts where ethical and practical, compare predeclared slices, and use targeted human review. I do not continuously train from positive clicks as if they were clean labels.

Google's production ML guidance warns about training-serving skew and feedback loops. An AI system adds more sources of skew because model, prompt, tools and retrieval can all change the observed distribution.

Turn evidence into cases safely

When production finds a meaningful failure, I do not copy the raw conversation into a permanent repository.

I first create a minimal case that preserves the failure mechanism:

production evidence
→ classify risk and root cause
→ remove or replace personal and confidential data
→ minimize the reproducer
→ obtain required review/consent
→ add expected contract and provenance
→ place in held-out or regression set

The production-trace workflow gives the privacy and retention details.

Not every poor rating becomes a regression. It may describe a product preference, missing capability, provider outage or unclear UI. I preserve the cause, not merely the unhappy output.

Use different alarms for different speeds

I split measurement by response time:

seconds: unauthorized calls, errors, latency, cost spikes
hours:   sampled quality, tool completion, fallback use
days:    repeated contact, reversions, retained task success
weeks:   user trust, adoption, workflow and business outcomes

Fast safety signals can stop a canary automatically. Slow product metrics should not be forced into an instant launch decision. I keep the previous complete configuration ready for rollback while delayed signals mature.

My practical rule

I never ask offline or online evaluation to prove what it cannot observe.

Offline cases create a repeatable release contract. A canary limits exposure. Production signals reveal distribution, operational and product reality. Version IDs connect the evidence. Human review interprets ambiguous cases. Sanitized failures improve the next offline suite.

This is a loop, but not a blind self-training loop. Each transfer between production evidence and permanent evaluation data is an engineering and privacy decision. That is how online signals correct the prediction without turning noisy user behaviour into automatic truth.