01
What Evidence an AI Feature Needs Before You Ship It
Ship readiness for an AI feature is one artifact per failure class, each with a stated minimum bar and a stated way it can be invalidated. Built here with a runnable 300 item fixture.
Evals, retrieval, tools, and observability
A software-engineering view of AI-enabled products: where models help, where deterministic boundaries remain necessary, and how to operate the whole system responsibly.
A useful AI feature is still a production system. It depends on permissions, identity, current data, tool contracts, latency budgets, evaluation, and a recovery path when a probabilistic component behaves differently than expected.
This series focuses on those engineering boundaries. The model is important, but it is one component inside a system that still needs clear invariants and evidence that it works.
01
Use models for useful ambiguity, while permissions, state transitions, money, and irreversible actions remain enforceable in code.
02
Measure retrieval, tool selection, outputs, latency, and cost against representative tasks instead of judging a demo by feel.
03
Record the inputs, versions, sources, tool calls, and decisions needed to explain and improve a production result.
Entry points into this subject, ordered as a reading path.
01
Ship readiness for an AI feature is one artifact per failure class, each with a stated minimum bar and a stated way it can be invalidated. Built here with a runnable 300 item fixture.
02
A trajectory grader for detecting unauthorized reads, duplicate effects, fabricated tool results, waste, and unsafe recovery even when an AI agent's final answer looks correct.
03
A simulation of a real 5 point improvement showing how often a fixed eval threshold gives a random red build, and why a paired comparison detects the gain at a much smaller sample size.
04
Agent retries must reuse one logical action identity and reconcile uncertain effects. An effect ledger makes autonomy bounded, inspectable, and recoverable.
Through the layers
A reference flow for keeping user identity, tenant, audience, and allowed actions attached to every AI tool call without trusting the model to enforce permissions.
Through the layers
Structured model output is not trusted application input. Validate syntax, semantics, authorization, and effect policy again at the tool execution boundary.
Measurement
A working record-and-replay harness for agent tests: a JSON cassette of decisions and tool results, deterministic replay with zero live calls, and a readable diff when the trajectory changes.
Measurement
A practical agreement experiment for measuring an LLM judge against human labels, including ambiguous cases, position bias, confidence, and release thresholds.
Measurement
A retrieval benchmark for separating missing evidence from generation failures using Recall@k, Precision@k, MRR, unanswerable questions, and controlled oracle context.
Measurement
A versioned evaluation matrix for separating prompt, model, retrieval, tool, and grader changes while protecting safety, quality, latency, and cost.
Measurement
A graded assertion approach for regression-testing variable AI outputs without comparing every generated word to one frozen answer.
Measurement
A measurement architecture that joins repeatable offline evals with production outcomes, safety signals, canaries, and privacy-aware review.
Measurement
A practical AI failure taxonomy that separates what happened, where it happened, its impact, and the suspected cause so teams can build useful metrics.
Measurement
A purely random sample hides rare and expensive AI failures. Combine baseline, risk-triggered, and exploratory samples without corrupting the denominator.
Measurement
A strong system prompt cannot secure retrieval, tools, authorization, retries, or side effects. Red-team the complete AI workflow and its recovery paths.
Measurement
A practical pipeline for learning from production AI failures while minimizing personal data, preserving the failure signal, and keeping evaluation fixtures reviewable.
Older, introductory or narrower pieces on the same subject.
Engineering context
I help teams turn AI-enabled workflows into maintainable products with clear data, evaluation, permission, and failure boundaries.