The Memory Hole Gets Bigger: How Corporate Websites Are Quietly Erasing the Past
The Memory Hole Gets Bigger: How Corporate Websites Are Erasing the Past
In 2019, General Electric quietly restructured its investor relations website. Old press releases — including several from the mid-2000s in which company executives made specific, verifiable promises about financial performance — were migrated to a new URL structure. Some made the trip. A lot didn't. The original URLs, which had been cited in news articles, academic papers, shareholder lawsuits, and regulatory filings, now return 404 errors.
Nobody announced this. Nobody archived it in time. And now those primary sources — the actual words of actual executives — exist only in quotes and summaries written by journalists who may or may not have transcribed them accurately. The original documents are just gone.
This is not an isolated incident. This is the web in 2024.
The Scale of the Problem Is Staggering
The Wayback Machine at the Internet Archive has been heroically crawling and preserving web content since 1996. It currently holds over 800 billion web pages. That sounds like a lot until you understand what it doesn't have: the vast majority of content that existed only briefly, was never indexed before deletion, existed behind login walls, or was stored in formats that crawlers couldn't capture.
A 2021 study by researchers at Harvard Law School found that approximately 25% of links in New York Times articles published between 1996 and 2019 were broken. For Supreme Court opinions — legal documents that are supposed to be the permanent record of American jurisprudence — the number was even higher. Links cited in judicial opinions as supporting evidence were dead at a rate that should alarm anyone who cares about the integrity of legal reasoning.
And that's just links that someone bothered to study. The broader web is rotting faster than any archiving effort can keep pace with.
The Corporate Redesign as Historical Erasure
When a company redesigns its website, the primary concern is almost never historical preservation. The concern is brand alignment, SEO performance, and whether the new color palette tests well with focus groups. The old URLs are collateral damage.
Sometimes there's a redirect strategy — old links get pointed to new locations. More often, there isn't. The content management system gets replaced, the URL structure changes, and thousands of links that existed in news articles, academic citations, and user bookmarks become dead ends overnight.
"The corporate website redesign is probably the single most destructive event for digital historical record," says one digital archivist who works with a major research university and asked not to be named because their institution has contracts with several of the companies they were criticizing. "Companies do it every three to five years, they almost never implement proper redirects, and nobody outside the organization even knows it's happening until the links are already dead."
The problem is compounded when companies get acquired. When a tech startup gets absorbed into a larger corporation, its website — which may contain years of blog posts, product documentation, press releases, and public statements — often gets redirected to the acquirer's homepage within weeks. The content itself is deleted. Not archived. Deleted.
News Organizations Are Not Innocent Here
It would be convenient to frame this as a corporate behavior problem distinct from journalism, but that would be dishonest.
News organizations have been deleting and restructuring content at an alarming rate, often for reasons that have nothing to do with accuracy or editorial judgment. URL restructuring during CMS migrations. Paywalling old free content without maintaining the original URLs. Deleting articles that attracted legal attention. Pruning "low-performing" content to improve SEO metrics — a practice that has become widespread enough to have its own industry euphemism: "content pruning."
When a regional newspaper decides that articles from 2009 aren't driving enough traffic to justify the server overhead and deletes them, it's not just losing page views. It's destroying the primary record of local events — city council votes, school board decisions, business openings and closings — that may exist nowhere else.
The Internet Archive has tried to fill this gap, but it can't be everywhere. Its crawlers visit most sites infrequently. Content that exists for a few days before deletion often goes uncaptured entirely.
The Fact-Checking Nightmare
Here's the practical consequence that affects ordinary people trying to understand the world: it is becoming progressively harder to verify claims made in the past.
A politician says they never supported a particular policy. The article in which they announced that policy is gone — the news site restructured, the URL is dead, the Wayback Machine captured the homepage but not the article. The quote exists in a tweet that references the article, but the tweet links to the dead URL, and the original article text is nowhere accessible.
A company claims it has always been committed to environmental responsibility. The press releases from 2008 in which executives argued against environmental regulations are gone — deleted during a rebranding exercise, never archived. The claims exist only in contemporaneous news coverage, which is itself increasingly behind paywalls or deleted.
This isn't a hypothetical. Researchers, journalists, and fact-checkers describe this situation in nearly every conversation about digital research methodology. The web was supposed to make it easier to verify claims against primary sources. Instead, it's created a landscape where primary sources have an expiration date set by whoever owns the server.
The Archivists Are Fighting With One Hand Tied
The Internet Archive is the closest thing the web has to an institutional memory, and it operates on a budget that would barely cover the catering at a mid-sized tech company's all-hands meeting. It has faced legal challenges from publishers who argue that archiving constitutes copyright infringement — a position that, if upheld broadly, would effectively make preservation illegal.
Individual archivists and volunteer communities do heroic work. The Archive Team, a loose collective of volunteer preservationists, has raced to capture content from platforms announcing shutdowns, often downloading millions of pages in the hours between a closure announcement and the actual server going dark. They saved significant portions of GeoCities, early Tumblr, and dozens of smaller platforms that would otherwise be completely gone.
But they can't be everywhere, and they can't capture what they don't know is disappearing. Corporate website migrations don't come with press releases.
What Gets Lost When History Disappears
The most important thing being lost isn't nostalgia. It's accountability.
The web created an unprecedented opportunity for permanent, accessible public record. Every public statement, every press release, every corporate announcement could in theory be preserved indefinitely and referenced by anyone. That opportunity is being squandered by routine infrastructure decisions made by people who have never thought about their implications.
When the links rot, the record rots with them. And unlike a physical archive that deteriorates slowly and visibly, digital deletion is instant and invisible. One day the page is there. The next day it isn't. And the day after that, most people have forgotten it ever existed.
The internet is broken. We take notes. Increasingly, though, nobody's taking notes about the notes that used to exist.