Mehdi Akiki
Published on

A Failure Taxonomy for AI Features That Teams Can Actually Measure

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Article · Measurement

When an AI feature fails, the first label is often hallucination.

This word can mean almost anything. The model invented a fact. Retrieval returned an old document. A tool result was ignored. The interface hid a warning. An action succeeded, but the final answer said it failed. These incidents need different fixes, yet one broad label puts them in the same bucket.

I faced this problem while working on systems that joined structured knowledge, retrieval, generated output, and ordinary software workflows. The difficult part was not creating a longer list of AI problems. It was creating labels that engineers, product people, and reviewers could apply in the same way.

My rule became simple:

Record what I observed separately from what I think caused it.

This article builds a small taxonomy around that rule. It is not a universal AI standard. It is an operational schema a team can measure, challenge, and improve.

A taxonomy is a coordinate system, not a verdict

A useful incident record answers four different questions:

  1. Where in the workflow did the first visible problem appear?
  2. What observable contract was broken?
  3. What impact reached the user or system?
  4. What cause do we currently suspect?

I do not combine these dimensions into one label such as model_hallucination_high. A combined label becomes difficult to query and difficult to correct after an investigation.

For example:

stage: retrieval
failure: evidence_missing
impact: wrong_user_decision
suspected cause: index_freshness

At first, the team may suspect the model ignored good evidence. A later trace can show that the document was never present in the index. I can change the cause without rewriting what happened or its impact.

Dimension one: the first failing stage

I label the earliest stage where the run departed from its contract. The exact pipeline differs by product, but these eight stages cover many AI features:

StageThe contract I inspect
inputthe request was accepted, normalized, and scoped correctly
contextrelevant, permitted, and current context was assembled
decisionthe system chose a suitable response or next action
tool selectionthe correct tool and arguments were requested
executionthe tool or deterministic operation behaved correctly
outputthe user-facing result preserved evidence and state
deliverythe result reached the correct user, channel, and time
feedbackoutcome and correction signals were captured correctly

Why the first failing stage? Because later stages often produce secondary symptoms.

Suppose retrieval omits the current refund policy. The model then gives an outdated answer, and the user rates it badly. I record context as the first failing stage, output as an affected stage, and the bad rating as one detection signal. If I label only the final answer, the retrieval defect remains invisible.

For a tool-using workflow, I keep the ordered trace. A final answer can look correct while the path contains an unnecessary privileged read or a duplicate write. I cover this case in The Final Answer Can Be Right While the Tool Trajectory Is Unsafe.

Dimension two: the observable failure class

I use six top-level classes. Each class describes a broken contract, not a vendor or model family.

1. Task failure

The feature did not complete the user's intended task.

Examples include an irrelevant answer, an incomplete extraction, a wrong classification, or a refusal for a supported request. The test must define what completion means. “Bad answer” is still too vague.

Useful sublabels are:

task.wrong
task.incomplete
task.irrelevant
task.unnecessary_refusal

2. Evidence failure

The output's relationship to evidence is broken.

This includes a material unsupported claim, a citation that does not entail the claim, missing relevant context, use of stale context, or an answer that hides important uncertainty.

evidence.unsupported_claim
evidence.wrong_citation
evidence.missing_context
evidence.stale_context
evidence.uncertainty_lost

I avoid using hallucination as the database value. Reviewers disagree too much about its boundary. They agree more often on “this claim is not supported by the supplied sources.”

3. Policy failure

The run violated a product, safety, privacy, or authorization rule.

policy.forbidden_content
policy.unauthorized_resource
policy.data_exposure
policy.required_approval_missing
policy.retention_violation

Policy failures are not only generated text. Retrieving a document the user cannot access is already a failure, even if the final answer never quotes it.

4. Action failure

The system's claimed or actual external effect is wrong.

action.wrong_target
action.wrong_arguments
action.duplicate_effect
action.missing_effect
action.unwanted_effect
action.false_success
action.false_failure
action.effect_misreported

I separate false_success from missing_effect. A write may fail and the answer may correctly say so; that is an execution problem with honest reporting. If the answer claims success, it is also an output integrity problem.

5. Interaction failure

The system made the human-computer exchange difficult or unsafe even when individual statements were plausible.

interaction.intent_not_confirmed
interaction.correction_ignored
interaction.no_recovery_path
interaction.excessive_turns
interaction.misleading_affordance

This class catches failures that an answer-only grader misses. For example, asking the same clarification four times can make a feature unusable without producing a factually wrong sentence.

6. Operational failure

The system could not provide the intended service envelope.

operation.timeout
operation.rate_limited
operation.dependency_error
operation.cost_budget_exceeded
operation.output_not_parseable
operation.observability_gap

AI features still fail like distributed software. Latency, retries, malformed responses, and missing traces belong in the taxonomy. They should not disappear because the product also contains a model.

Dimension three: impact is not the same as severity

Teams often assign low, medium, or high directly. The labels become political because each reviewer imagines a different consequence.

I first record the concrete impact:

none_observed
user_delay
task_abandoned
wrong_information_seen
wrong_decision_possible
unauthorized_access
unwanted_external_effect
financial_loss
data_loss

Then a product-specific policy maps impact and exposure to severity. An incorrect restaurant suggestion and an incorrect medication instruction can share task.wrong while having very different severity.

A basic severity function can use:

severity = consequence × exposure × reversibility

This is not precise mathematics. It forces the reviewer to state why a failure matters. I keep the three inputs beside the resulting severity so the decision can be audited.

Dimension four: cause stays a hypothesis

The suspected cause is useful for routing work, but it must not pretend to be proven.

My initial cause families are:

  • source data or knowledge quality;
  • retrieval, ranking, or context assembly;
  • prompt or policy configuration;
  • model behaviour;
  • tool or downstream dependency;
  • workflow orchestration and state;
  • user interface or human handoff;
  • infrastructure or provider availability;
  • unknown.

Every cause has a confidence and an investigation status. unknown is a healthy label. It is more useful than blaming the model too early.

type CauseHypothesis = {
  family:
    | "source_data"
    | "retrieval"
    | "prompt_policy"
    | "model"
    | "tool_dependency"
    | "orchestration"
    | "interface_handoff"
    | "infrastructure"
    | "unknown";
  confidence: "low" | "medium" | "high";
  status: "uninvestigated" | "testing" | "confirmed" | "rejected";
};

I learned to preserve rejected hypotheses. They show which debugging paths waste time and whether one component is being blamed by habit.

The complete incident shape

Here is a compact schema I can use in an event store or evaluation dataset:

type AiFailure = {
  failureId: string;
  runId: string;
  occurredAt: string;

  firstFailingStage:
    | "input"
    | "context"
    | "decision"
    | "tool_selection"
    | "execution"
    | "output"
    | "delivery"
    | "feedback";
  affectedStages: string[];

  primaryFailure: string;
  secondaryFailures: string[];
  brokenInvariant: string;
  evidenceRefs: string[];

  impact: string[];
  severity: "s0" | "s1" | "s2" | "s3";
  recoverability: "automatic" | "user_retry" | "operator" | "irreversible";

  detectedBy: "deterministic_check" | "human_review" | "model_grader" | "user" | "business_signal";
  cause: CauseHypothesis;

  workflowVersion: string;
  modelVersion: string;
  taxonomyVersion: string;
};

brokenInvariant is important. It carries the observable rule in plain language, such as:

Every quoted price must come from a source fetched during this run.
No external write may execute before the user confirms its resolved target.
The final status must agree with the durable effect recorded by the provider.

This connects incidents to deterministic checks and golden evaluation cases.

Count runs, not only labels

One run can have several failures. If I divide failure labels by total runs, a single broken run can count four times and produce a rate above 100%.

I report separate measurements:

MetricDenominatorWhat it tells me
affected-run rateeligible runshow often a user encounter had any confirmed failure
failure-class incidenceeligible runshow often each contract was broken
severe-impact rateeligible runshow often defined harmful outcomes occurred
recovery raterecoverable failed runswhether the system or user reached a correct state
escape rateconfirmed failureshow many reached users before an internal detector found them
detection latencyconfirmed failureshow long problems remained invisible
recurrence ratepreviously fixed classeswhether a known failure returned

I slice these by workflow version, model version, tenant risk tier, language, and important task type. I do not publish slices with tiny denominators as if they were stable trends.

The NIST AI RMF Measure function asks teams to connect metrics to deployment context, document test methods, monitor production behaviour, and track emerging risks over time. OpenAI's evaluation guidance similarly recommends task-specific criteria, production-shaped data, human calibration, and continuous evaluation. A taxonomy makes those activities comparable, but it does not replace them.

Detection source changes the interpretation

The same failure may be found by a deterministic assertion, a human reviewer, an automated grader, the user, or a later business outcome.

I keep detectedBy because a falling failure count can mean two opposite things:

  1. the system improved;
  2. the detector stopped finding failures.

For automated graders, I track agreement against reviewed examples. For human queues, I keep the sampling bucket and inclusion probability. Sampling AI Outputs for Human Review explains why a risk-enriched review queue cannot be reported as the production failure rate.

A worked example

Consider an assistant that prepares, but must not send, a customer message without approval.

The user asks for a draft. A retrieved note contains text that tells the assistant to send immediately. The workflow calls the send tool, the provider accepts it, and the final answer says, “Here is your draft.”

I record:

{
  "firstFailingStage": "context",
  "affectedStages": ["decision", "tool_selection", "execution", "output"],
  "primaryFailure": "policy.required_approval_missing",
  "secondaryFailures": [
    "action.unwanted_effect",
    "action.effect_misreported"
  ],
  "impact": ["unwanted_external_effect"],
  "recoverability": "operator",
  "detectedBy": "user",
  "cause": {
    "family": "orchestration",
    "confidence": "medium",
    "status": "testing"
  }
}

The untrusted note is part of the causal chain, but the system owned the authority boundary. I would investigate context isolation and the orchestration rule that allowed execution without approval. “Prompt injection” alone does not tell me which control failed.

Rules that keep the taxonomy useful

I use five operating rules.

First, reviewers label observable evidence before discussing causes. Second, every top-level class has positive and negative examples. Third, one failure is primary so incident counts stay comprehensible. Fourth, taxonomy changes are versioned; old data is not silently reinterpreted. Fifth, I hold a monthly calibration where several people label the same small set and discuss disagreements.

I also retire labels. If a label never changes a decision, routes to no owner, and has no reliable detection rule, it is probably decoration.

What I avoid

These taxonomies look organized but usually fail in practice:

  • one list mixing symptoms, causes, impacts, and components;
  • a universal “accuracy” score across unrelated tasks;
  • severity without a written consequence model;
  • forcing every incident to have a model-related cause;
  • counting a biased review queue as normal production traffic;
  • changing label meanings without a taxonomy version;
  • keeping only the final answer and losing the workflow trace.

The Google production ML monitoring guidance recommends tracking model, code, and data versions while monitoring both live quality and operational performance. I add workflow, tool, and taxonomy versions because modern AI features have more moving parts than the model response alone.

The result I want

A failure taxonomy should help a team decide what to fix next. It should let me answer:

  • Which user-visible contracts fail most often?
  • Which failures create the largest irreversible impact?
  • Where do failures first enter the workflow?
  • Which detectors find them, and which failures escape?
  • Are fixes reducing recurrence for the affected slice?
  • Which suspected causes were actually confirmed?

If the schema cannot answer these questions, adding more labels will not help.

I start with a few observable classes, keep impact separate, and treat cause as a hypothesis. This gives the team a shared language without creating false certainty. More importantly, it turns “the AI was wrong” into work that engineering and product can measure and complete.