Investigation
This site as an engineering laboratory
I started with one question: why is a mostly static site this heavy? I kept going until my intuition was not reliable any more. Several of the most promising ideas lost, a few problems turned out to live in the framework and the browser, and I stopped when six experiments planned in advance all failed their own gates.
How to read this page
Every investigation below separates four things, because they are easy to mix up and the mix-up is where most performance claims go wrong.
- Result: what changed on this site, or that nothing changed.
- Experiment: the protocol, scripts and raw data. Early work only kept single traces, and it is labelled that way.
- Conclusion: the decision rule I took away from it.
- Upstream: something found below my application, in Next.js, Chrome DevTools or a library, with my exact role and whether it is submitted.
All timings are lab measurements. The analytics on the live site does not collect performance data yet, so no result here is a field result. Lab-only conclusions are marked.
The original question
On September 7, 2026 the homepage of a mostly static site transferred 829,434 bytes in 17 requests. A custom webpack setting forced every dependency into one shared vendor chunk. The full search index, 2.48 MB raw, was fetched at startup even if nobody opened search. The Rust hub and the Atlas landing page serialized hundreds of entries into single documents of 1.45 MB and 3.29 MB. And the site shipped Web Vitals code that reported to an analytics sink that did not exist.
So most of the cost was not in the pages. It was global work that every route paid for. That is why the first changes were about ownership, not about compressing assets.
How I measured
- Production builds served locally with
next start, never the development server. - A fixed Chromium mobile profile: 390 × 844 viewport, Fast 4G, 4× CPU slowdown. Bytes come from Resource Timing, not from build reports.
- Later experiments have their own folder with scripts and retained data. Where timing matters, runs alternate between variants in fresh browser contexts.
- The last experiments, PERF-050 to PERF-056, have protocols with go and no-go gates that were written before the run.
- A faster result is not accepted if it changes behavior. Screenshots, keyboard paths, no-JavaScript paths and back navigation are checked alongside the timings.
- Failed and excluded runs are kept and labelled, not deleted.
Before and after, for context
These numbers matter because of what they showed, not as a score. They are lab byte counts, which are deterministic for a given build.
- Homepage: 17 requests and 829,434 transferred bytes, down to 13 requests and 146,265 bytes (82.4% less) after the first four changes.
- Search index at startup: 2.48 MB raw, down to zero bytes until someone opens search.
- Atlas landing page: gzip HTML 91.9% smaller and DOM 90.3% smaller, with every case still reachable without JavaScript.
- An isolated Tailwind compile: 77.2 seconds and 14.0 GiB of memory, down to 355 ms and 157 MiB.
The layers it went through
The questions started in application code and kept moving down when the application could not explain the measurement:
- Application configuration and the bundler's chunking rules.
- React server and client boundaries, and the Flight payload they produce.
- The Next.js router's prefetch protocol and its task scheduler.
- HTML delivery: compressed prefix size, head order, HTTP revalidation.
- The browser image pipeline: priority, decode, raster, presentation.
- CSS delivery as a cache graph across a journey, not as isolated files.
- Tailwind's compiler inputs, and the cost of producing the site.
- Chrome DevTools' trace model, when a warning blamed the wrong thread.
- HTTP compression dictionaries (RFC 9842) and font fallback metrics.
Investigations
Grouped by mechanism, not by date. PERF numbers refer to the experiment log; the folders hold the protocols and data.
Who owns the work before first paint?
PERF-001, 003, 004, 005, 011, 020
- Result
- A custom vendor chunk put every dependency on every route, the full search index loaded at startup, and analytics started before the page finished. Moving optional work off the default path cut the homepage from 829,434 to 146,265 transferred bytes after the first four changes.
- Conclusion
- Look for global work before polishing assets. One attempt lost: calling dynamic() from a Server Component did not keep the footer form out of startup, so the boundary had to move instead.
- Experiment
- Single historical measurement: byte counts are deterministic, timings come from one trace, and the raw trace was not kept.
Idle prefetch is a bandwidth policy
PERF-002, 018
- Result
- Viewport prefetch fetched route data for links nobody followed. With intent prefetch, /blog, /notes and /open-source went from 7, 8 and 3 idle fetches to zero, and a hover-then-click still reached the page in 50 to 58 ms.
- Conclusion
- Prefetching everything visible is a decision about other people's bandwidth. Prefetch on intent unless the data says otherwise.
- Experiment
- Single historical measurement: byte counts are deterministic, timings come from one trace, and the raw trace was not kept.
Search: the slow part was not the network
PERF-022, 023
- Result
- The search index shipped the authoring model. A four-field protocol made it 86.6% smaller. Then loading everything in parallel made the input appear later (391 to 643 ms), because 941 commands were initialized before it painted. A small eager shell with an intent-only index won.
- Conclusion
- The faster network path can still lose to initialization work. Measure until the input is usable, not until the bytes arrive.
- Experiment
- Single historical measurement: byte counts are deterministic, timings come from one trace, and the raw trace was not kept.
A crawlable archive without shipping the archive
PERF-006, 007, 008, 009
- Result
- Large hub pages serialized hundreds of entries into one document. Bounded directories cut the Atlas landing page's gzip HTML by 91.9% and its DOM by 90.3%, and the /rust document by 92%. Every link stays reachable without JavaScript.
- Conclusion
- A smaller document is not automatically a faster paint: one single trace even got 64 ms slower, so no paint improvement is claimed for that step.
- Experiment
- Single historical measurement: byte counts are deterministic, timings come from one trace, and the raw trace was not kept.
When the Server Component is the bigger choice
PERF-015, 021, 024, 025, 026
- Result
- Moving the header to a Server Component saved 99 bytes of JavaScript but made the document and the Flight payload larger, so it was reverted. Eight small client links lost to one delegated island. The official multiple-root layout still loaded other routes' chunks.
- Conclusion
- Evaluate a server boundary where it changes the serialization graph, not only by the JavaScript it removes.
- Experiment
next-rendering-boundary. PERF-015 has a folder with summary medians but no raw rows. The rest are single historical measurements.
The smaller editor was slower
PERF-019
- Result
- Three curated Monaco builds, about 21% smaller than the CDN version, became usable later than the full local build. Evaluation and grammar setup, not bytes, decided the ready time. The full local build was kept until the editor was removed from the site in October 2026.
- Conclusion
- Bytes are a proxy. When the proxy and the user-visible milestone disagree, trust the milestone.
- Experiment
- Single historical measurement: byte counts are deterministic, timings come from one trace, and the raw trace was not kept.
Upgrading without letting the new version erase the baseline
PERF-016, 017, 028
- Result
- Turbopack repeated more Flight chunk-list bytes on the homepage (2.05% against a 2% budget) than webpack (0.95%), so production stayed on webpack behind an executable budget. DevTools estimated 12.1 kB of legacy JavaScript; the removable amount was 298 bytes.
- Conclusion
- Treat the bundler as a variable and turn a regression report into a release gate that fails the build.
- Experiment
next-legacy-polyfills. PERF-016 and 017 are single historical measurements plus the budget script in scripts/.- Upstream
- Comment draft on Next.js #86785 / #88551 (unsubmitted).
What does the framework cost, and what breaks without it?
PERF-029, 030
- Result
- Plain HTML improved cold LCP by only 17 to 19 ms, but roughly halved load time. A static tier with 89% fewer startup bytes still lost warm navigation, because the live router reuses the DOM it already has. The framework stayed.
- Conclusion
- There is no universally fastest rendering model. Cold entry, warm navigation and back navigation each have a different winner.
- Experiment
next-content-delivery-floor,static-content-tier- Upstream
- Validation for Next.js Route Handler compression, issue #98007 / PR #98044 (unsubmitted).
When does prerender pay for itself?
PERF-031 to 035
- Result
- Speculation Rules prerender helped only after real intent. Immediate touch and focus triggers were wasted on abandonment. Waiting for prefetch to finish before prerendering was slower, up to 632% at one boundary, because Chrome could already overlap the work.
- Conclusion
- Speculation needs an abandonment budget and proof of activation. The intuitive sequential pipeline was the slow one.
- Experiment
static-content-tier
The first 2 KiB of HTML is a scheduling API
PERF-036, 037, 038
- Result
- In the static-tier prototype, 2 KiB was the smallest compressed prefix that worked across both transports. Reordering the head recovered 60 ms without adding bytes. Content-addressed assets removed six 304 revalidations.
- Conclusion
- Order and cache policy schedule work as much as size does. A 304 is still a round trip.
- Experiment
static-content-tier
Image eligibility, priority and preload
PERF-039 to 043
- Result
- The first MDX diagram stays lazy but starts at high priority, and desktop gets a media-gated preload that moved LCP 532 to 924 ms earlier. A two-candidate preload duplicated downloads and was 500 to 812 ms slower at fractional pixel ratios.
- Conclusion
- Priority and eagerness are different decisions. Preload is an exception to lazy loading, not a default.
- Experiment
mdx-image-priority,mdx-media-preload,mdx-preload-metadata,next-image-sizes-parser,next-image-preload-media- Upstream
- Original Next.js patches: a sizes parser that misses calc() and decimals, and a preloadMedia option for Image (both unsubmitted).
The image finished, so why has it not painted?
PERF-044, 045
- Result
- The time after the image arrived was layout, raster and presentation. Synchronous decode and an early flush both lost. Separately, seven visible tags triggered 11 idle route requests after LCP; reusing intent links removed about 131 kB. An earlier theme script was rejected because it caused a hydration repair and a duplicate image download.
- Conclusion
- Attribute the whole interval before blaming the network or JavaScript. Work after the headline metric still costs.
The CSS frontier: faster-looking options that lost
PERF-048, 049, 051, 052
- Result
- Inlining all CSS won 372 to 480 ms of cold LCP and was still rejected: 18 to 22 kB more cold transfer, mixed or slower reloads, and 836 MB more build output. Splitting the historical article stylesheet by rendered content was kept. Critical CSS and route-weighted splitting failed their gates.
- Conclusion
- A better headline metric is not a better system when the full journey, the cache and the build all get worse.
- Upstream
- Independent evidence for Next.js issue #95141 on inline CSS duplication (not posted).
Tailwind scanned 12 GB of old builds
PERF-050
- Result
- Automatic source detection was reading old build directories. Bounding the sources cut an isolated compile from 77.2 s and 14.0 GiB to 355 ms and 157 MiB, and the stylesheet lost 883 gzip bytes. Page load did not change measurably.
- Conclusion
- Some performance work is about the cost of producing the site, not the cost of loading it.
- Experiment
tailwind-source-boundary
When the fifth prefetch disappeared
PERF-046, 047
- Result
- Some prefetches never ran. Decoding the router's prefetch protocol led to a Next.js scheduler bug where sibling tasks were lost. It was already reported. I reproduced it and confirmed the existing fix red/green with both webpack and Turbopack. No speed claim survived the measurement.
- Conclusion
- A correctness fix does not need a latency story. Report the bug, not an invented speedup.
- Experiment
next-prefetch-sibling-starvation- Upstream
- Validation of Next.js issue #96965 / PR #97377 (review note not posted).
Making the lab hard to fool
PERF-010, 013, 027
- Result
- Web Vitals code was reporting to an analytics sink that did not exist, an audit passed without testing anything, and a DevTools forced-reflow warning came from another thread. They became a removal, an executable check, and a DevTools patch draft.
- Conclusion
- Instruments can be confidently wrong. Check the instrument before trusting the reading.
- Experiment
devtools-forced-reflow-attribution. PERF-010 and 013 are single historical measurements.- Upstream
- Original Chrome DevTools patch with a test (unsubmitted).
The wall: six planned experiments, stopped by their own gates
After the first fifty experiments I chose the next candidates in advance and wrote the gates before running them. All six stopped.
| Experiment | Gate set in advance | What happened |
|---|---|---|
Cacheable critical CSS core with the full stylesheet deferred | Exact rendering, at most 4 KiB gzip blocking CSS, at most 2 KiB added cold transfer | The small arms broke the page during rapid scroll. Every correct arm duplicated too many bytes. |
Route-weighted CSS split | At least 20% or 3 KiB smaller route-weighted CSS, 10% better journey transfer, Flight share under 2% | 7.55% worse weighted bytes, 10.9% worse journey, 3.83% Flight share. |
HTTP 103 Early Hints for the critical stylesheet PERF-053 | Needed a stable critical asset from PERF-051 | Never eligible to run, because PERF-051 produced no asset worth hinting. |
Compression dictionary transport for article pages | 25% fewer document bytes by page three, dictionary cost repaid by page three | Primed pages were 59% smaller, but fetching the dictionary made a three-page visit 11.8% (64 KiB) and 55.6% (128 KiB) worse. |
Prerender only popular pages, render the tail on demand | 50% less server output, 30% faster build, 20% less memory | 29% less output, 5% faster, 0.4% less memory. |
Tuned fallback font metrics for the cold font swap | No route gets worse | Every metric arm made at least one route worse. Hiding or withholding the font to fake a zero shift was refused. |
Why I stopped there
I did not stop because I ran out of ideas. There is still a queue. I stopped because each remaining idea could only win under conditions I cannot show exist yet: reading sessions longer than three pages for dictionaries, an edge path with room for Early Hints, a smaller shared CSS core, or field data showing that a cold font shift of at most 0.0023 matters.
The gates were written before the runs, so I could not move them after seeing the numbers. And the CSS work had already shown the trap: inlining all CSS improved cold LCP by 372 to 480 ms, and it was still the worse system once reloads, the full journey and the build were measured. Shipping a change like that only because one metric looks better is the decision I wanted to avoid.
Upstream
Some problems were not in my code. The roles below are exact: finding a bug is different from validating a fix someone else already wrote. None of these is submitted yet.
| Finding | Layer | My role | Status |
|---|---|---|---|
| Prefetch scheduler loses sibling tasks | Next.js | Reproduced and validated the existing fix (issue #96965, PR #97377) red/green with webpack and Turbopack | Unsubmitted review note |
| Route Handler responses skip compression | Next.js | Validated the existing fix (issue #98007, PR #98044) on real content | Unsubmitted comment |
| Image sizes parser misses calc() and decimals | Next.js | Original finding and patch | Unsubmitted, not yet run in the Next.js monorepo |
| Image preload cannot carry a media query | Next.js | Original patch for a long-requested gap (discussions #25171, #29621, #71393) | Unsubmitted |
| Webpack ships all module polyfills | Next.js | Webpack parity evidence for issue #86785 / PR #88551 | Unsubmitted comment |
| Inline CSS duplicated into every RSC payload | Next.js | Independent evidence for existing issue #95141 (+836 MB build output) | Not posted |
| Forced reflow blamed on the wrong thread | Chrome DevTools | Original finding, patch and test | Unsubmitted, not yet run in a DevTools checkout |
| Umami script cannot load after the page | Pliny | Original finding and patch | Unsubmitted |
What is not here yet
- Field data. Every timing above is from the lab until real-user LCP, INP and CLS are collected.
- Raw data for the early experiments. Some of them will be measured again against the version of the site from before this work.
- RSC-aware edge caching (PERF-012) was planned but never run, so there is no result to report.
- Separate write-ups for each investigation. They will be linked from here.