Mehdi Akiki
Published on

Sampling AI Outputs for Human Review Without Only Seeing Easy Cases

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Article · Measurement

Human review sounds simple: take some AI outputs every day and ask people whether they are good.

The first dashboard usually looks reassuring. Most outputs are ordinary, short, and easy. Reviewers see many correct answers. Meanwhile, the rare long conversation, unusual language, dangerous tool call, or expensive customer workflow receives almost no attention.

I learned to separate two questions:

  1. How often does the system fail in normal traffic?
  2. Where can the system fail badly, even if that traffic is rare?

One sample cannot answer both questions well.

Random sampling is necessary, but not sufficient

A uniform random sample gives every eligible output the same probability of selection. This is useful for estimating an overall rate.

If 2% of reviewed random outputs contain a defined failure, I can estimate the production failure rate with uncertainty around it. The sample has a real denominator.

But suppose only 0.1% of traffic can trigger a high-impact action. In a daily sample of 200, I will often see none of those cases. A clean review tells me little about the risky path.

The mistake is not random sampling. The mistake is asking it to discover every important failure.

Begin with the sampling frame

Before choosing records, I define the complete set from which they may be selected. I call this the sampling frame.

For an AI workflow, one row may represent:

  • one final response;
  • one conversation;
  • one tool trajectory;
  • one document extraction;
  • one autonomous task run.

These units are not interchangeable. Sampling responses over-represents long conversations. Sampling conversations can hide repeated bad steps inside one run. For an agent, I usually keep both a task-level record and its ordered tool-call events.

I also record the eligibility rule. For example:

completed production tasks
excluding internal tests
with consented, reviewable content
between 00:00 and 23:59 UTC

Without this definition, the denominator changes silently.

Use three complementary buckets

My practical review queue has three buckets.

BucketPurposeTypical share
baseline randomestimate ordinary production quality50%
risk-triggeredinspect known high-consequence paths35%
explorationsearch underrepresented and novel cases15%

The percentages are examples, not universal defaults. The important point is to label the bucket on every reviewed item.

Baseline random sample

This bucket is selected uniformly from the full eligible population. It protects the measurement from my assumptions. New common failures can appear here even when no detector knows about them.

Risk-triggered sample

This bucket deliberately over-samples events such as:

  • a write or delete tool was requested;
  • policy filters disagreed;
  • retrieval returned low evidence coverage;
  • a response was unusually long or expensive;
  • the user retried, corrected, or abandoned the workflow;
  • the model changed its answer after a tool error;
  • input or output language is rare in normal reviews.

A trigger is not proof of failure. It is a reason to spend review capacity.

Exploration sample

Known triggers only find known shapes. I reserve capacity for novelty: rare feature combinations, new tenants, new models, new prompt versions, outliers in embedding or telemetry space, and small traffic segments that have not been reviewed recently.

This bucket prevents the risk rules from becoming a map of yesterday's incidents.

Stratify before the queue becomes convenient

If reviewers can choose the next easy item, they will naturally reduce queue discomfort. That introduces selection bias.

I assign strata before review. Useful dimensions include:

product workflow × risk tier × language × model version × outcome class

I do not create every possible combination. Sparse strata become impossible to operate. I choose dimensions tied to a concrete hypothesis, then enforce a minimum review count for important small groups.

For example, a translation feature may be 80% English-to-French and 1% English-to-Arabic. Uniform traffic sampling will be dominated by French. A small minimum for Arabic helps detect a segment failure, while the baseline sample still estimates overall traffic.

Keep inclusion probability with every item

Risk sampling changes the apparent failure rate. If tool-writing tasks are 1% of traffic but 30% of the review queue, the raw review percentage is not a production percentage.

I store why a record was selected and, when I need an aggregate estimate, its inclusion probability:

type ReviewCandidate = {
  runId: string;
  bucket: "baseline" | "risk" | "exploration";
  stratum: string;
  inclusionProbability: number | null;
  triggerCodes: string[];
};

The baseline bucket can support a straightforward rate estimate. A stratified sample can be reweighted if selection probabilities are known. An exploration sample often has no honest population estimate, and I label it as discovery evidence instead of forcing it into one number.

This small metadata field prevents a serious dashboard lie.

Do not sample only by model confidence

Model confidence, a judge score, or a safety classifier can be useful triggers. They should not control the entire sample.

The model and detector can share the same blind spot. A confidently wrong answer will not enter a “low confidence” queue. A judge trained on similar preferences may approve the same weak pattern.

I use automated signals to enrich the risk bucket, then keep independent random and exploration buckets.

Make the review question observable

“Is this response good?” creates inconsistent labels. I ask narrower questions tied to evidence:

  • Does every material claim have support in the supplied context?
  • Did the workflow request only tools needed for the task?
  • Was user authorization checked before the external action?
  • Did the answer preserve uncertainty from the source?
  • Could a reasonable user recover after the failure?

For each label, I add examples and a short decision rule. I periodically send the same item to two reviewers. Disagreement is not reviewer noise to hide; it shows that the rubric or the product expectation is unclear.

The NIST AI Risk Management Framework recommends representative evaluation, documented test methods, and production monitoring. Its Generative AI Profile also treats human moderation and active discovery of unexpected outputs as parts of risk management. I translate that into a review system with explicit populations, selection rules, and follow-up actions.

Close the loop without training on everything

A reviewed failure can lead to different actions:

product bug      → fix deterministic code
missing evidence → improve retrieval or refuse
unsafe action    → reduce permissions or add approval
model behaviour  → prompt, model, or policy evaluation
unclear rubric   → clarify the product contract

I first preserve the failing example as a regression case when policy permits. I do not immediately place every reviewed conversation into a training dataset. Production records can contain private data, temporary context, reviewer mistakes, and distribution bias from the sampling strategy.

Review data needs purpose, retention, access control, and provenance like any other sensitive dataset.

Metrics that remain honest

I keep at least four views:

  1. Baseline estimated failure rate, with sample size and uncertainty.
  2. Failure rate by predeclared stratum.
  3. Trigger yield: how often each risk rule finds a confirmed issue.
  4. Discovery log: new failure classes found by exploration or reviewers.

Trigger yield helps remove useless rules. A rule that selects thousands of harmless outputs consumes reviewer time. A rare rule that repeatedly finds high-impact failures may deserve more capacity even if it does not affect the overall percentage much.

I also watch time-to-review. A perfect sample reviewed three weeks later may be too slow for a bad deployment.

A small operating rhythm

For a new production workflow, I start with this rhythm:

  • daily baseline and risk review;
  • weekly exploration of underrepresented segments;
  • weekly disagreement calibration between reviewers;
  • a review after every model, prompt, retrieval, or tool-policy change;
  • monthly retirement or adjustment of sampling triggers.

The exact numbers depend on volume and consequence. The structure is what matters.

What I carry into production

Human review is not a pile of examples. It is a measurement system and a discovery system sharing one queue.

I preserve a random baseline for honest rates. I over-sample known risks to use scarce attention well. I keep an exploration budget for failures I have not named yet. Most importantly, I never combine these buckets into one impressive but meaningless quality score.

That is how review finds the hard cases without losing sight of normal production reality.