Photo by James sojan on Unsplash
Somewhere on the internet right now, there is a news article citing a study. The study is linked. The link goes to a page that says '404 — Page Not Found.' The study existed. The link existed. The relationship between them — the chain of evidence that makes a citation meaningful — no longer exists.
This is happening millions of times a day. We have a name for it: link rot. We don't have a solution for it. We don't even have a particularly good way to measure how bad it's gotten, which is itself a symptom of how bad it's gotten.
The Scale of the Problem Is Staggering
In 2021, Harvard Law School's Perma.cc project published research finding that more than 70% of URLs cited in Supreme Court opinions returned errors or had fundamentally changed their content. Not blog posts. Not random forum threads. Supreme Court opinions — legal documents that are supposed to be permanent records of reasoning and precedent, citing sources that have quietly vanished.
A separate study found that roughly half of all URLs cited in academic papers published between 1997 and 2012 were dead within a decade. The New York Times, ProPublica, and other major outlets have done their own internal audits and found that significant percentages of their own archived articles contain broken outbound links — sources that were accurate when cited and have since evaporated.
The web was designed around the assumption that links would persist. That assumption was wrong, and we've spent thirty years building an information infrastructure on top of it.
Why Content Disappears: The Business Logic of Deletion
Content doesn't vanish from the internet because of entropy. It vanishes because someone made a decision — or failed to make one.
The most common culprit is the website redesign. A company rebuilds its site, migrates to a new CMS, restructures its URL taxonomy, and fails to implement redirects for the old addresses. Every link pointing to the old URLs now goes nowhere. This isn't malicious. It's negligence — the kind of negligence that's nearly invisible to the people doing it, because they're looking at the shiny new site, not at the thousands of incoming links they just severed.
Then there's the death of publications. The internet graveyard of shuttered digital media outlets is enormous — Deadspin (original version), The Awl, Pacific Standard, Splinter, Gawker, countless local news sites that were acquired, gutted, and eventually deleted. When these outlets disappear, every article they ever published — and every outbound link in those articles, and every inbound link pointing at them from other sources — becomes garbage data.
Government websites are particularly brutal offenders. Federal and state agencies regularly restructure their web presence without preserving old URLs, destroying links that journalists, researchers, and citizens spent years building into documents and databases. The EPA, USDA, and various other agencies have quietly made enormous amounts of public information unreachable simply by moving it without redirects.
And then there's the deliberate deletion — the article that gets pulled for legal reasons, the press release that becomes inconvenient, the product page for software that's been discontinued, the study that got retracted. These disappearances are intentional, and they leave a specific kind of damage: the citation that still points somewhere, but now points to a blank page or a generic homepage, with no indication that anything was ever there.
The Citation That Proves Nothing
For journalism, link rot creates a verification crisis that's easy to underestimate.
One of the core functions of hyperlinking in online journalism was supposed to be radical transparency — you could follow every claim back to its source, evaluate the evidence yourself, check the reporter's work. This was genuinely new. Print journalism cited sources but rarely let you examine them directly. The web made the evidence chain walkable.
Except now, for a growing percentage of articles, the chain is broken. You can follow the link. You'll find a 404. You have no way to evaluate whether the original source actually said what the article claimed, whether it's been updated or retracted since the article was written, or whether it ever existed at all. The citation becomes a gesture toward evidence rather than evidence itself.
This matters more than it might seem in an era of contested facts and misinformation. When bad actors want to launder fabricated claims into the information ecosystem, one effective technique is to cite sources that no longer exist. The claim looks sourced. The source can't be checked. The 404 does the work that the nonexistent evidence can't.
The Wayback Machine Is Not Enough
The Internet Archive's Wayback Machine is genuinely one of the most important cultural institutions on the internet — an underfunded, under-appreciated nonprofit that has been crawling and archiving the web since 1996 and currently holds over 800 billion web pages. If you want to find a dead link, the Wayback Machine is often your best shot.
But it's not a solution. It's a heroic workaround being performed by an organization that runs on donations and has been the target of multiple significant cyberattacks, including a breach in 2024 that exposed user data and temporarily took the service offline. The entire web's institutional memory is being maintained by one nonprofit with an uncertain future, and we've collectively decided that's fine.
It's not fine. The Wayback Machine doesn't capture everything — it misses paywalled content, dynamically loaded pages, content that robots.txt files excluded from crawling, and the enormous volume of content that was deleted before it could be archived. It also doesn't integrate with the web in any meaningful way — finding a dead link's archived version requires a manual process that most readers, journalists, and researchers simply don't perform.
The Incentive Structure That Keeps Nothing Fixed
Link rot persists because the people who create it don't bear the costs of it.
When a company redesigns its website and breaks ten thousand inbound links, the company pays nothing. The journalists, researchers, and ordinary users who relied on those links absorb all the costs — in time spent searching for the moved content, in citations that become unprovable, in knowledge that becomes inaccessible. The externalities are entirely offloaded.
Until that changes — until there are meaningful incentives for content creators and platform operators to maintain the links they publish — the web's trash problem will keep growing. Every new article published today is a future 404 in waiting, a future broken citation, a future piece of evidence that will one day prove nothing because it points to nothing.
The internet was supposed to be a permanent, searchable record of human knowledge. Instead, it turns out to be more like a whiteboard that anyone can erase at any time, with no obligation to tell you what was written there before.
We've built an information infrastructure on sand, and the tide comes in a little further every year.