Let the Bots Roam: Why We Should Preserve Digital Decay
The common cry in digital preservation is one of perfect capture. We build web crawlers to be meticulous, respectful, and thorough. We instruct them to honor robots.txt files, to tread lightly on servers, to be the polite guests of the web. Our goal is a pristine snapshot, a complete record of a page as it stood at a precise moment. But what if this politeness is costing us something vital? What if, in our quest for clean data, we are sanitizing history itself?
I want to make a counterintuitive argument: we should sometimes let preservation be messy, intrusive, and even flawed. We should allow—or even design—crawlers that capture the friction of the live web. The 404 error page, cleverly designed by a sysadmin, is a cultural artifact. The "bandwidth exceeded" message from a forgotten GeoCities site is part of its story. The sluggish load time of a server on its last legs, the broken image hotlinked from a now-dead domain, the comment from a scraping bot that itself is now extinct—these are not noise. They are signal.
The Aesthetics of Failure as Historical Record
When we curate only the successful fetch, we present a digital past that never truly existed. The live web was, and is, a cacophony of failure states. It stutters, breaks, and resists. A perfectly preserved HTML file, stripped of its original loading context, tells a lie of seamless accessibility. It erases the texture of technological limitation, the struggle of bandwidth, and the reality of entropy in real-time.
Think of it as preserving not just the painting, but the cracked varnish and the fading pigment. A medieval manuscript shows wormholes and water stains; these physical traces tell a story of its journey through time. Our digital equivalents are the server errors, the ad-blocker voids, the CAPTCHAs, and the cookie warnings. A crawler that blindly stumbles through these, capturing them as part of the page’s "state," is creating a more honest record. It captures the web as it was *experienced*, not just as it was ideally rendered.
This isn't a call for disrespect or denial-of-service attacks. It’s a philosophical shift. Perhaps alongside our "polite" archival crawls, we need designated "feral" ones. Their mandate wouldn't be completeness, but phenomenology. They would run with scripts enabled, on emulated period-specific browser engines, collecting the full sensory and interactive failure of a page. They would hit rate limits and record the consequences. They would, in essence, archive the strain.
In sterilizing our captures, we risk building archives of ghosts—clean, quiet, and utterly divorced from the chaotic, frustrating, and beautiful reality of the networked world. To truly preserve the digital experience, we must make room for the glitch, the lag, and the dead end. We must allow the bots to sometimes roam, not as careful librarians, but as witnesses to the decay that is an inseparable part of the life they seek to record.
Notes & further reading
A few pages I came back to while writing this:
- Elk Grove, CA
- The Broken Chain of St. Gall: How a Medieval Monk Preserved What a Digital Age Might Lose
- Pasadena, CA
- The Flickering of a Dead Man's Hand
- New Haven, CT
- The Memory of a Single String: How a Tiny URL Holds a Universe
- Stamford, CT
- Washington, DC
- one area's overview
- a practical rundown
- Little Rock, AR
- Gilbert, AZ
- Peoria, AZ