The Conflicted Archivist: To Hoard the Web or To Curate It
In the quiet, humming rooms that house our collective digital memory, a quiet philosophical war has been waged for decades. It isn't fought with weapons, but with spiders—web crawlers, to be precise. On one side stands the principle of comprehensive capture; on the other, the practice of curated selection. These are the two opposing poles of web archiving, and the space between them defines what future generations will know of our present.
The comprehensive approach, famously championed by organizations like the Internet Archive, operates on a breathtaking scale. Its mission is one of benevolent hoarding: to capture everything, from the most trafficked news site to the most obscure personal blog. The guiding ethos is that it is not our place to judge what may have value later. A forgotten GeoCities page about a teenager's pet iguana might one day be a priceless artifact for a sociologist studying early internet culture. This method is automated, relentless, and democratic. It seeks to create a complete, if messy, fossil record of the digital age, trusting that context and meaning will be applied by researchers in the distant future.
In stark opposition stands the curated model, employed by many national libraries and academic institutions. This approach is not automated but mindful. Archivists, acting as digital librarians, make conscious decisions about which websites, online publications, or social media streams to preserve. They select materials based on thematic collections, perceived cultural significance, or legal deposit requirements. The goal here is not quantity but qualified quality. It is the difference between saving every book from a publishing house and saving only the ones that win literary awards. This method creates a leaner, more navigable archive, but one that is inherently filtered through the biases and blind spots of its human curators.
The tension between these methods is the central drama of digital preservation. The comprehensive crawl risks drowning truly significant material in an ocean of digital noise, making it virtually impossible to find. Yet, the curated selection risks committing acts of premature epistemic violence, silencing voices and entire communities deemed unworthy by contemporary standards. What if the curator overlooks a nascent political movement organizing on a fringe forum? What cultural treasures are lost because they weren't trending on the day the archivist pressed 'record'?
There is no perfect answer, only a necessary conflict. The health of our historical record may depend on this very tension. We need the massive, undiscriminating crawl to serve as a safety net, capturing the raw, unvarnished chaos of the web. And we need the careful, curated collection to provide a focused, intelligible narrative. One gathers the universe of data; the other draws the constellations. Our future understanding will likely be a dialogue between the two—a constant questioning of what was saved, why it was saved, and what, in the silent gaps between them, was allowed to disappear forever.
Notes & further reading
A few pages I came back to while writing this:
- Dayton, OH
- The Unassuming Archive of the Shopping List
- Toledo, OH
- The Midwinter Seed Vault: Dormant Data and the Promise of Spring
- Oklahoma City, OK
- The Fantasy of the Perpetual Custodian: Why No Archive Is an Island
- Tulsa, OK
- Eugene, OR
- Portland, OR
- Salem, OR
- Philadelphia, PA
- Pittsburgh, PA
- Charleston, SC