This site as an engineering laboratory
Kept, rejected and upstream results from more than fifty experiments on this site.
Evidence: experiments/
Treating an assumption as a hypothesis, building an experiment, following a surprising result into a lower layer, and rejecting a change when the evidence says it loses.
Kept, rejected and upstream results from more than fifty experiments on this site.
Evidence: experiments/
3 parts
A measured diagnostic tree for slow Rust builds: cargo --timings for the shape, the fingerprint log for unexpected rebuilds, -Ztime-passes for the phase inside one crate, and what splitting a workspace into more crates actually changes.
Evidence: experiments/rust-atlas/build-times/
Ship readiness for an AI feature is one artifact per failure class, each with a stated minimum bar and a stated way it can be invalidated. Built here with a runnable 300 item fixture.
A simulation of a real 5 point improvement showing how often a fixed eval threshold gives a random red build, and why a paired comparison detects the gain at a much smaller sample size.
A working record-and-replay harness for agent tests: a JSON cassette of decisions and tool results, deterministic replay with zero live calls, and a readable diff when the trajectory changes.
A trajectory grader for detecting unauthorized reads, duplicate effects, fabricated tool results, waste, and unsafe recovery even when an AI agent's final answer looks correct.
A practical agreement experiment for measuring an LLM judge against human labels, including ambiguous cases, position bias, confidence, and release thresholds.
A retrieval benchmark for separating missing evidence from generation failures using Recall@k, Precision@k, MRR, unanswerable questions, and controlled oracle context.
A versioned evaluation matrix for separating prompt, model, retrieval, tool, and grader changes while protecting safety, quality, latency, and cost.
A graded assertion approach for regression-testing variable AI outputs without comparing every generated word to one frozen answer.
A measurement architecture that joins repeatable offline evals with production outcomes, safety signals, canaries, and privacy-aware review.
A practical AI failure taxonomy that separates what happened, where it happened, its impact, and the suspected cause so teams can build useful metrics.
A purely random sample hides rare and expensive AI failures. Combine baseline, risk-triggered, and exploratory samples without corrupting the denominator.
A strong system prompt cannot secure retrieval, tools, authorization, retries, or side effects. Red-team the complete AI workflow and its recovery paths.
A practical pipeline for learning from production AI failures while minimizing personal data, preserving the failure signal, and keeping evaluation fixtures reviewable.