Measurement
This site as an engineering laboratory
Kept, rejected and upstream results from more than fifty experiments on this site.
Evidence: experiments/
The main body of this writing follows one system below its interface: a compiler, a runtime, a protocol, a build, a data pipeline. It is organized around four engineering questions.
Investigations run over several parts and keep the programs and experiments that produced their evidence. Articles stand alone. Shorter notes, beginner material and tool guides stay in the archive and in search.
Multi-part work backed by code, measurements or reproductions that can be run again.
Measurement
Kept, rejected and upstream results from more than fifty experiments on this site.
Evidence: experiments/
Through the layers · 10 parts
Ten parts on what a type is and where it goes: sets of values, layouts and padding, erasure and monomorphization, the bits the processor actually sees, and what the checker knows that the binary forgets. Small programs in Rust, C, TypeScript and Python, with one script that reproduces every output.
Evidence: experiments/rust-atlas/types-under-the-hood/
Derived state · 3 parts
A measured diagnostic tree for slow Rust builds: cargo --timings for the shape, the fingerprint log for unexpected rebuilds, -Ztime-passes for the phase inside one crate, and what splitting a workspace into more crates actually changes.
Evidence: experiments/rust-atlas/build-times/
Interrupted execution · 3 parts
Five async Rust problems reproduced in one small tokio program: an oversized future, a future that is not Send, a blocking call, a select! loop that loses data, and a shutdown that hangs. Then the cost of async fn through dyn Trait, and cleanup when Drop cannot await.
Evidence: experiments/rust-atlas/async-symptoms/
Through the layers
Rust failures that are difficult to name and easy to misdiagnose, organized from symptom to mechanism, with a failing and a repaired fixture for most cases.
Standalone pieces, grouped by the question they answer.
Through the layers
Source, runtime, protocol, compiler, machine. Each piece follows one behaviour down to the layer that decides it.
Investigations: Types under the hood, Async Rust under the hood, Rust Failure Atlas
One small Rust function followed through real HIR, THIR, and MIR output, with a practical explanation of what each compiler representation is for.
Most WebSocket tutorials start at the API. This one starts at the wire — electrons, NIC interrupts, kernel socket buffers, the TCP state machine, the HTTP upgrade handshake, the binary frame format — and works up to what ws.send() actually does.
A practitioner’s deep dive into ops, resources, and the V8 bridge in Deno’s Rust core.
Zero-copy scanning still reads every inspected byte. It removes a second materialization in memory. I use streaming Aho-Corasick to explain automaton state, chunk boundaries, DFA dependency chains, prefilters, and honest performance measurement.
Two identifiers, two different jobs. DefId names a definition across the whole compiler pipeline and across crates. HirId points to a node in the syntax tree of the current crate. Mixing them up is the first stumbling block when reading rustc source code.
Derived state
What gets stored, what gets recomputed, what becomes stale, and how the system knows.
Investigations: Rust build times
Rust's incremental cache is a dependency graph, not a saved compiler process. Learn what an edit dirties, how green nodes are recovered, and why Cargo may still run rustc.
A practical model of rustc queries, providers, memoization, dependency tracking, cycles, and red-green incremental reuse.
When you sync two systems with webhooks, the copy slowly becomes wrong: some events are lost, some are duplicated, some arrive in the wrong order. Reconciliation is the way to make the two systems agree again. Here is how to design an API for it.
Cargo build scripts are cached from declared file and environment inputs. Use rerun-if rules, verbose evidence, and a rebuild matrix to remove accidental rebuilds.
Concrete simulations of the records a timestamp cursor can miss, plus safer tuple cursors, overlap windows, change logs, and reconciliation.
Interrupted execution
Crashes, retries, cancellation, partial progress, ownership, recovery and idempotency.
Investigations: Async Rust under the hood
A checkpoint is a claim about durable effects, not the last page fetched. Compare commit protocols and crash points so recovery repeats work instead of losing it.
A paginated import should resume from a committed cursor and stable item identities, not from a human page number that may no longer describe the same data.
Async cancellation usually means dropping a future after Pending. See how select loops lose partial progress and how to design restartable operations.
A concurrent refresh design using local single-flight work, database versions, conditional writes, and recovery rules for providers that rotate refresh tokens.
Agent retries must reuse one logical action identity and reconcile uncertain effects. An effect ledger makes autonomy bounded, inspectable, and recoverable.
Measurement
Treating an assumption as a hypothesis, building an experiment, following a surprising result into a lower layer, and rejecting a change when the evidence says it loses.
Investigations: This site as an engineering laboratory, Rust build times
Ship readiness for an AI feature is one artifact per failure class, each with a stated minimum bar and a stated way it can be invalidated. Built here with a runnable 300 item fixture.
A simulation of a real 5 point improvement showing how often a fixed eval threshold gives a random red build, and why a paired comparison detects the gain at a much smaller sample size.
A working record-and-replay harness for agent tests: a JSON cassette of decisions and tool results, deterministic replay with zero live calls, and a readable diff when the trajectory changes.
A trajectory grader for detecting unauthorized reads, duplicate effects, fabricated tool results, waste, and unsafe recovery even when an AI agent's final answer looks correct.
A practical agreement experiment for measuring an LLM judge against human labels, including ambiguous cases, position bias, confidence, and release thresholds.
All 267 published pieces, newest first, each labelled as an investigation, an article or reference. Reference covers beginner Rust, framework and tool guides, AI explainers, interview preparation and older short pieces. The 61 short engineering notes have their own page.