
The Knowledge Half-Life
It's 9 AM in New Jersey and a deployment fails. The error is a schema mismatch: the integration expects a retry_count field in the response. That field was renamed to attempt_number four days ago when the payments team refactored the retry logic, and nobody updated the Notion page because nobody remembered they had to. The current deployment still works. It's not urgent. But the engineer wants to fix it before the next release.
So the ritual begins. Re-ask the engineer who owns the payments API, except she's in San Francisco where it's 6 AM. Read the actual response payloads in the logs and figure out what the current schema looks like. Forty minutes gone before the actual fix, a one-line field rename, gets committed.
This ritual happens every sprint, every incident review, every planning session where someone has to figure out what the notification pipeline actually does before anyone can make a decision about changing it. The team treats it as diligence. It's waste. And nobody has named what causes it, because the Confluence diagram with three boxes still labeled "TBD" doesn't look like a failure mode. It looks like normal engineering.
The re-verification loop isn't a habit. It's a tax, and the rate depends on variables you haven't measured yet.
Margaret Storey, a computer science professor at the University of Victoria who studies how developers actually work, gave it a name in February 2026. As she put it, technical debt lives in the code, but cognitive debt lives in developers' minds. That phrasing, and it landed in software engineering because it names the gap precisely. Peter Naur, the Naur in Backus-Naur form, the notation behind every programming language grammar, made the same observation in 1985: that a program is a theory living in the minds of its developers, and the code is a lossy representation of that theory. The term "cognitive debt" was circulating in adjacent domains since mid-2025, but Storey's articulation landed for engineering specifically.
Her follow-up post from February 18 listed what cognitive debt looks like from the inside: loss of confidence making changes, heavier review burden, debugging friction, slower onboarding, low-grade stress and fatigue that teams have come to treat as baseline. If you've managed an engineering team for more than eighteen months, you've seen these. You've probably attributed them to growth pains or the wrong people rather than to a structural erosion that happens in every codebase regardless of who's on the team.
"Velocity can outpace understanding." That's Storey's phrasing, and it's the quiet version of the stale API docs.
Having a name for this matters practically, not just conceptually. Once you've named cognitive debt as a distinct phenomenon, separate from technical debt and with its own accumulation mechanics, you can start asking what drives the accumulation rate. And once you can ask that question, you can intervene on specific causes instead of throwing process at a symptom you've never characterized.
Technical debt is visible. SonarQube gives you a score, the linter fires, the code review flags the smell. Cognitive debt has none of this instrumentation, and the asymmetry matters more than it sounds.
There is no alert that fires when the person everyone Slacks with questions no longer actually understands the code, and no dashboard shows the gap between what the architecture diagram says and what the system is doing in production. The decay is happening continuously. You only find out about it when a deployment fails on a field that was quietly renamed days earlier and forty minutes vanish before a one-line fix gets committed.
The "why" behind a decision is the fastest-decaying component. Hermann Ebbinghaus, the psychologist who first measured how memory fades, established over a century ago that we forget roughly 70% of new information within 24 hours, and that was for rote memorization of meaningless syllables under controlled lab conditions. Contextual intent decays faster, because memory reconstruction is schema-driven: your brain fills in what "should have happened" based on what you know now, systematically erasing the anomalies and edge cases that drove the actual decision. The harder the original reasoning was to compress, the worse the reconstruction. Research on after-action reviews consistently finds that immediate debriefs produce more accurate and usable findings than delayed ones. The wiki page written six weeks after the architecture meeting is not the architecture meeting. It's a story about the architecture meeting, with the hard parts smoothed over and the contradictions resolved in retrospect.
A 2010 study from the University of Zurich built a computational model for developer knowledge that adds a second mechanism of erosion: it's not just time that erodes your understanding, it's the system changing around you. When others commit to a file you wrote, your knowledge score decreases, measurably rather than metaphorically. Unlike the time-continuous decay Ebbinghaus documented, this erosion is event-driven: every sprint where someone else modifies a component while your attention is elsewhere, that specific component's knowledge score degrades. We covered this in Post 4 on ownership drift. The connection is direct: both forms of decay are running simultaneously, and nobody's watching either rate.
Standard exponential decay describes this precisely. If K(t) is the institutional knowledge your team holds about a system at time t, and K₀ is the knowledge at the moment of capture, then:
K(t) = K₀e^{-λt}
(I'm aware I've just introduced a differential equation into a newsletter about engineering management. The useful part is not the formula. It's what λ measures and what controls it.)
K(t) and K₀ are not things you can directly measure. But λ, the decay constant, is driven by factors your team touches every week. High turnover accelerates λ: every departure takes knowledge that was never fully externalized, and the reconstruction at the next hire is never complete. Rapid hiring that dilutes the ratio of context-holders to newcomers accelerates λ. Lack of structured capture at peak recall accelerates λ: the decision gets made, someone sets a calendar reminder to document it, the reminder gets snoozed, and the understanding decays before it's ever written down.
The 9 AM deployment failure, IMO, is what three of these accelerators stacking on the same system looks like. The API owner is three time zones behind and hasn't started her day, which means no overlap before the gap became operational. The schema changed during a refactor days earlier, and whatever reasoning lived in her head when she renamed that field was never captured while she still had it. The docs weren't updated, and the API contract stayed accurate at authorship only to become a liability in production. None of that is negligence. It's just decay running at an unmanaged rate on a system nobody knew to watch.
What drives λ down? Capture at peak recall, meaning within hours of a decision or an incident, not six weeks later. Deliberate re-engagement when underlying systems change, meaning a trigger, not a recurring task, that brings documentation current when the thing it documents changes. Structured overlap between the person who holds the theory and the person who will need to understand it, before the knowledge gap becomes operational.
Research suggests sector-level variation in these rates. Medical knowledge carries a documented half-life around two years (Harvard Medical School, 2017). Software engineering benchmarks cluster around 12 to 18 months. In the AI age, where the effective best practice reshuffles every few months, practitioners report the working half-life of specific techniques is shorter still, though no one has put a clean number on it. If your wiki is refreshed annually, you're very likely operating on significantly degraded knowledge for your core systems, not as a possibility but as a mathematical consequence of the decay rate.
What AI changes is the rate: more code gets generated, the same understanding bandwidth gets diluted, and the gap between what's in the codebase and what's in people's heads grows faster than any documentation cadence can realistically keep up with. AI doesn't introduce cognitive debt as a new phenomenon. It raises λ, which is a different and more specific problem.
Lambda is not a fixed property of a team. It's driven by behaviors: how quickly decisions get documented after they're made, how consistently documentation gets refreshed when the underlying systems change, how much deliberate overlap exists between the people who hold knowledge and the people who will need it.
The interrupt model for knowledge transfer, which Post 3 in this series established as broken, is broken partly because of timing. It moves knowledge reactively, after the gap becomes painful, at which point the source knowledge has itself been decaying. The ownership labels Post 4 showed you don't track the current state of expertise. They track the state at authorship, and λ is why. The labels were probably accurate once, and time passing and systems changing are what make them unreliable now.
The question isn't whether knowledge decays. It decays on every team, in every codebase, at every company. The question is whether anyone is watching the rate.
Pick your most critical service. Ask someone who's been on the team for a year but hasn't touched that service in six months to explain why the retry logic works the way it does. If they can't answer without pinging the person who built it, that's your λ, and it's already higher than you think.
Next time: If knowledge decays this predictably, can you detect where it's concentrated before it's gone, extract the reasoning while it's still fresh, and ground it somewhere your team can actually query? That's the loop.