The Immortal Data Lie: We Are Not Preserving the Web, We Are Taxidermying It

There’s a quiet faith that underpins our work in web archiving, a received wisdom we rarely question: that by capturing the HTML, the CSS, the images, we are preserving a webpage. We point to the Wayback Machine with a sort of reverent pride, a digital Noah’s Ark saving bits from the flood of link rot. But I’ve come to believe this is a comforting illusion. What we are doing is not preservation; it’s taxidermy. We are stuffing the carcass of a webpage and propping its glassy-eyed form in a static diorama, then calling it a living record.

Consider the simple hyperlink, the web’s foundational gesture. In a live page, a link is a promise, a potential, a doorway. Clicking it is an act of navigation, of discovery, of context. In an archived page, a link is a tragic artifact. It points, with brittle hope, to a snapshot that may or may not exist elsewhere in the archive, or it points into the void. The essential quality of the link—its dynamism, its interconnectedness—is the first thing to die in the archival process. We preserve the shell of the connection but euthanize its function, like preserving a handshake by severing the arm and mounting it on a plaque.

This taxidermy extends to the lived experience of the web. We might capture the text of a bustling forum from 2005, but we cannot capture the feeling of waiting for a dial-up connection to load it, the specific anxiety of a page rendering line by line, the camaraderie of a real-time conversation that unfolded over hours. We preserve the pixels of a Flash animation but not the shared cultural moment of everyone on the internet simultaneously watching a dancing baby or playing a crude browser game. The software emulators needed to run these relics are our equivalent of the formaldehyde jar, a desperate attempt to simulate a spark of life in a long-dead medium.

The Context That Slips Through the Net

The most profound loss is one of performance and periphery. A captured page is silent regarding its own loading speed, which dictated user patience and engagement. It tells us nothing of the browser wars that shaped its rendering, the plugins it demanded, the screen resolutions for which it was designed. The social context evaporates: the unspoken rules of a now-defunct social platform, the inside jokes that scroll away into oblivion, the ambient pressure of the algorithmic feed that governed what you saw and when. We save the tweet, but we lose the timeline.

This is not to say the effort is worthless. A taxidermied specimen still holds immense value for scientists; it shows morphology, scale, color. Similarly, an archived webpage provides crucial evidence of what was, a bulwark against revisionism and forgetting. But we must be honest about the profound limitations of our craft. We are not building a living library. We are creating a hall of fossils.

Admitting this changes the mission. It makes us more humble and perhaps more strategic. It pushes us to document not just the page, but the *experience* of the page. To write down the stories of its use, to record the peripheral data streams, to acknowledge that some digital lifeforms are simply too ephemeral to stuff and mount. Perhaps the real preservation work is not in the robotic capture of billions of individual pages, but in the careful, human curation of the stories they told when they were still alive. The web was a conversation, and we have mostly succeeded in preserving the monologue. The echoes are gone.

Notes & further reading

A few pages I came back to while writing this: