Published on

How AI Is Eating Your Codebase From the Inside

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

A 30-second prompt generates 500 lines of code. A proper review takes 45 minutes. Nobody consistently picks the 45-minute option. And that asymmetry is slowly hollowing out engineering teams from the inside.

The contamination curve

Here's what I've seen happen, in teams I've worked with and in my own workflow after spending months with Claude Code.

The first time AI generates a chunk of code, you review it carefully. You read every line, trace the logic, check for subtle bugs. Good.

The second time, you review it again. Mostly. You skim the parts that look familiar.

The third time, you review less. The pattern feels safe now.

By the fifth time, you're reading the diff title and checking if CI is green. Tests pass? Merge.

You didn't decide to stop reviewing. It just... happened. And that's the thing nobody talks about: this isn't a discipline problem. It's a neuroscience problem.

Your brain is not on your side

Human brains are wired to minimize cognitive effort for a given result. This isn't laziness, it's how biological computation works. Your prefrontal cortex runs on glucose and optimizes ruthlessly for energy conservation. When you have two paths, "spend 45 minutes understanding this diff" versus "tests pass, CI is green, merge," your brain has to actively override the default every single time.

Willpower is a depleting resource across a workday. By the fourth PR of the afternoon, nobody is choosing the hard path. Not you. Not your tech lead. Not the senior engineer who's been shipping for fifteen years. The effort ratio is simply too absurd, and the feedback loop on skipped reviews is delayed: the bugs, the tech debt, the architectural drift show up weeks or months later, long after the dopamine hit of closing that ticket.

The one-way ratchet

Here's where it gets truly insidious: each skipped review doesn't just miss one batch of changes. It increases the cost of the next review.

Because now you're reviewing against a codebase you understand a little less. The gap between "what I know about this code" and "what this code actually does" widens with every cycle. And once AI-generated code that nobody fully understands is in the codebase, it becomes the context for the next round of AI generation. The AI builds on top of its own previous output, and your mental model falls further behind with each iteration.

This is a one-way ratchet. The divergence between prompting effort and review effort isn't linear, it's exponential. At some point the codebase becomes effectively authorless. No human on the team can explain why it's structured the way it is.

The green checkmark illusion

"But the tests pass!"

Right. And that's the most dangerous part, because it gives you a legitimate justification to skip the cognitive work. You're not being irresponsible. You have a genuine signal that says "this is probably fine."

The problem is that test suites verify behavior, not intent. They tell you the code does what it does, not whether what it does is what you actually wanted. Not whether the architecture makes sense. Not whether there's a subtle security issue. Not whether you've introduced a dependency that will bite you in six months.

But the green checkmark feels like understanding. It's a cognitive shortcut that looks like diligence.

The invisible metric

And this creates a perverse organizational dynamic.

The team's velocity metrics go through the roof. Management sees more PRs merged, more features shipped, more story points closed. By every visible measure, productivity has exploded.

The invisible measure of how well anyone on the team actually understands the system they're building has no dashboard. No metric. No sprint review slide. So the degradation is completely illegible to the organization until something breaks badly enough to force an investigation, and then you discover that nobody can explain how the service works anymore.

The darkest timeline

Here's the version that keeps me up at night.

Production goes down. You paste the logs into Claude. Claude suggests a fix. Tests pass. You deploy. Incident resolved.

The postmortem says "root cause identified and fixed" with a straight face. But nobody on the team actually understood the root cause. Nobody understood the fix. You've resolved an incident in a system you don't understand using a tool you can't verify. And next time it'll be even easier to do the same thing, because now you understand the system even less.

The incident postmortem is supposed to be the immune system of an engineering org. If even that gets delegated to AI, there's no feedback mechanism left.

What I've learned from watching this up close

I spent time onboarding a product team to Claude Code earlier this year. My contrarian takeaway, one that surprised me, was that AI coding tools surface coordination and discipline problems more than they accelerate writing code.

The teams that had strong review culture before AI kept it. They used AI to generate candidates and then applied the same scrutiny they always did. The tools made them faster at the generation step without degrading the verification step.

The teams that were already rubber-stamping PRs? They just found a way to rubber-stamp ten times faster.

The tool isn't the disease. It's an accelerant that reveals how thin the review culture already was.

The asymmetry nobody is pricing in

Here's the fundamental problem, stated plainly:

Prompting is cheap. Understanding is expensive. And the ratio gets worse over time.

Every cycle of "prompt -> generate -> skip review -> merge" makes the next review more expensive, which makes it more likely to be skipped, which makes the following review even more expensive. The cost of understanding grows while the cost of generating stays flat.

At some point, and I think a lot of teams are either at this point or approaching it fast, the codebase crosses a threshold where no amount of human effort can recover a full mental model. The only entity that "understands" the code is the AI that wrote it. And it doesn't actually understand anything.

So what do we do?

I don't have a clean answer. But I have some intuitions:

The people who will be most valuable in the next few years aren't the fastest prompters. They're the ones who can read, who can take an authorless codebase they've never seen before and tell you what it actually does, where the risks are, and what architectural decisions were made (or avoided). Deep code reading is becoming a rare and extremely valuable skill precisely because every incentive pushes against developing it.

At the organizational level, the teams that survive this will be the ones that treat code comprehension as a first-class metric, not just velocity, not just test coverage, but something like "can any two engineers on this team independently explain how this service handles [critical path X]?" If the answer is no, your velocity numbers are a fantasy.

And at the individual level? I think the only honest move is to be aware of the ratchet. Name it. Notice when you're three PRs deep and haven't actually read a diff. The awareness doesn't solve the problem, but it at least gives you a chance to intervene before the gap becomes unrecoverable.

The ratchet only turns one way. The question is whether you notice it turning.

I build and scale reliable production systems. Open to full-time and freelance work with U.S.-based teams that value ownership and execution.

Got something in mind?

Book a Discovery Call