Click It and Weep: The Quiet Disappearance of the Web You Thought You Saved
Somewhere in your browser, there's a bookmark folder you haven't opened in three years. Maybe it's called "Resources" or "Read Later" or just a sad little stack of unnamed folders nesting inside each other like Russian dolls of abandoned good intentions.
Go ahead. Open it. Click something.
We'll wait.
If your experience matches the data, roughly 25 percent of those links are already dead. Not redirected. Not archived. Just gone — a 404 page where something useful used to live.
This is digital rot, and it's eating the internet alive.
The Numbers Are Worse Than You Think
In 2023, the Pew Research Center published a study on link decay that should have caused a small panic in every newsroom, university, and government agency in the country. It didn't, because the findings were depressing in a slow, boring way rather than a dramatic one.
The study found that 38 percent of web pages that existed in 2013 are no longer accessible today. Among links cited in news articles, one in four lead nowhere. Supreme Court opinions — the kind of documents you'd assume someone, somewhere, is maintaining — contained broken links at a rate that made researchers visibly uncomfortable. Academic citations, government reports, medical literature: all riddled with URLs that resolve to error pages or domain squatters selling "premium web hosting solutions."
The average lifespan of a web page, depending on how you measure it, is somewhere between two and four years. The average lifespan of a printed book is several centuries. We have somehow built a civilization-scale information system that is less durable than a paperback.
Deletion as Strategy
Some link death is accidental — a server migration gone wrong, a CMS update that shuffled URLs, a startup that ran out of money and let the domain lapse. This is bad. It's also the charitable interpretation.
The less charitable interpretation is that a lot of content disappears on purpose.
Companies routinely scrub embarrassing press releases, product announcements for things that didn't work out, and blog posts that aged poorly. The practice is so common it barely registers as news. A tech company pivots, and suddenly the entire blog category for its previous business model returns a 404. A media outlet gets acquired, and the new owners quietly delete the archives rather than pay to host content they didn't produce. A government agency updates its policy position, and the old policy page — the one a thousand news articles linked to — vanishes without a redirect.
This isn't accidental. Redirects are not technically difficult. A competent web developer can implement them in an afternoon. When an organization chooses not to set up redirects during a URL restructure, that's a choice. And the thing that choice erases is the connective tissue of the web — the citations, the references, the trail of breadcrumbs that lets you verify where information came from.
The Wayback Machine Is Not Enough
The Internet Archive exists, and it is genuinely heroic. The nonprofit has been crawling and preserving the web since 1996, and its Wayback Machine has saved billions of pages that would otherwise be completely unrecoverable.
But it has limits, and those limits matter.
The Archive can't crawl everything. It especially struggles with content behind logins, paywalls, and JavaScript-heavy interfaces — which describes an increasing percentage of the modern web. It archives snapshots, not live databases, so dynamic content like comment sections, forums, and interactive tools often doesn't survive. And critically, it depends on pages being publicly accessible long enough to be crawled in the first place. Content that gets deleted within hours of publication may never be captured.
There's also the legal dimension. In 2024, a federal court ruled against the Internet Archive in a copyright case brought by major publishers, dealing a significant blow to its digital lending program. The message from the content industry was clear: preservation is fine, but not if it interferes with monetization.
Who Profits From Forgetting
It would be convenient if digital rot were purely a technical problem — a side effect of the web's chaotic growth and the inevitable entropy of complex systems. Some of it is. But some of it is a feature.
Consider the media industry. Publications that have been acquired, merged, or folded don't always survive in any meaningful archival sense. Gawker's archives — years of influential web culture coverage — went offline when the site was shut down following the Hulk Hogan lawsuit. Deadspin's history was disrupted through ownership changes. DNAinfo and Gothamist were shut down and briefly went dark entirely after staff voted to unionize. In each case, the historical record of those publications became a casualty of a business dispute.
Or consider the tech industry, where documentation is treated as a living document right up until the product is discontinued, at which point it becomes an embarrassing reminder that the product existed at all. Google has killed dozens of products — Reader, Stadia, Inbox, Allo, the list is genuinely long — and in many cases, the support documentation, community forums, and developer resources disappear with them. Developers who built integrations, wrote tutorials, or published guides based on those products are left with broken links and obsolete knowledge.
The beneficiary of all this forgetting is always the present. A company without a searchable history is a company that can reinvent itself without accountability. A platform without archives is a platform that controls its own narrative.
The False Horizon of Now
The practical effect of link decay is subtler than a missing page. It creates a false temporal horizon — a sense that the internet only meaningfully exists in the last six months or so.
If you're a student trying to research how a policy debate evolved, a journalist trying to verify a claim, or a developer trying to understand why a technical decision was made, the broken-link graveyard makes your job significantly harder. The absence of old content doesn't just make research inconvenient — it makes it unreliable. You can't cite what you can't access. You can't verify what's been deleted.
And into that vacuum rushes the present: new takes, new rewrites, new content designed to rank for the same queries the old content used to answer. The SEO content mills we mentioned earlier? They thrive on link decay. Every time a definitive old resource disappears, there's a gap in search results waiting to be filled by something newer and considerably worse.
The internet is not a library. Libraries have preservation mandates, trained archivists, and a professional ethic built around making information permanently accessible. The internet has terms of service, quarterly earnings calls, and a delete button.
So save that article you care about. Screenshot it. Print it, even. Because the link you're counting on being there tomorrow has already started its countdown.
The web doesn't remember. It just performs memory until it doesn't.