The Archive and the Firehose: Two Philosophies for Saving the Web
When we talk about preserving the vast, sprawling entity that is the web, two distinct archetypes emerge. They represent not just different technical approaches, but fundamentally opposing philosophies about what preservation means. On one side, there is the curated archive. On the other, the comprehensive firehose. One is a library; the other is a warehouse.
The curated archive is best exemplified by institutions like the Internet Archive. Its approach is humanistic and selective. It decides what is historically or culturally significant enough to save. It builds collections, creates thematic crawls, and presents its holdings with context. This is preservation as an act of interpretation. The archivist is a gardener, carefully tending specific plots, weeding out the irrelevant, and presenting a coherent landscape for future researchers. The value isn't just in the data itself, but in the story the collection tells.
In stark contrast stands the firehose approach, championed by projects like Common Crawl. This method is agnostic, automated, and massive. It doesn't discriminate. It aims to capture as much of the web as possible, as often as possible, and dump it all into storage. There is no curation, only collection. The goal is raw, exhaustive coverage—to save everything and let future generations figure out what it means. The archivist here is an engineer, maintaining the pipeline that siphons the entire ocean, one bucket at a time.
Each philosophy has its perils. The curated archive risks introducing the biases of its creators. What gets deemed 'unworthy' of saving vanishes forever, a silent censorship of the mundane. The firehose, however, risks creating an unusable mountain of data. Without curation, context is lost. A petabyte of raw HTML is a desert of information—it exists, but without a map or a guide, it's incredibly difficult to find meaning, let alone a single drop of water.
Ultimately, we need both. The firehose provides the raw, unvarnished bulk of our digital existence, a complete fossil record for a future that may have questions we cannot yet conceive. The curated archive provides the narrative, the focused lens that turns that raw data into understandable history. One gives us the entire forest; the other points out the most significant trees. In the grand project of saving our present for the future, the meticulous librarian and the industrial engineer are not rivals, but essential partners.
Notes & further reading
A few pages I came back to while writing this:
- Virginia Beach, VA
- The Library's Call Number: A Paper Ghost in the Machine
- Bellevue, WA
- The Digital Dust Bunny: What Your Browser Cache Knows About You
- Kent, WA
- The Myth of the Forever Database: Why Perpetual Access is a Dangerous Promise
- Spokane, WA
- Tacoma, WA
- Vancouver, WA
- Madison, WI
- Milwaukee, WI
- a useful directory
- a local resource