The Granary and the River: Confronting Preservation's Scale with Web Archives
There’s a fundamental tension at the heart of web archiving, one that pits two contrasting philosophies of preservation against each other. It’s not a battle of good versus evil, but rather a conflict of scale, priority, and ultimately, philosophy. We can think of these approaches as the Granary and the River. Most people, when they think of archives, picture the Granary: a stable, curated repository where specific, valuable things are selected, cataloged, and stored for the long term. The other, more overwhelming method, is the River: a continuous, automated attempt to capture the entire flow of the web, with all its chaos and ephemerality. The choice between building a granary and trying to map a river defines what we save and what we lose.
The Granary approach is human-scale. It’s what institutions like libraries have done for centuries. A librarian, or a team of archivists, makes conscious decisions. They identify a website, a collection of posts, or a set of public records deemed significant. They create a detailed catalog entry, perhaps capture the site at pivotal moments, and ensure the resulting files are stored in a preservable format. The value here is in selectivity and context. The resulting archive is high-fidelity and meaningful, but its scale is minuscule compared to the web itself. It’s like saving a single, perfect sheaf of wheat from a vast field; precious, but not representative of the entire harvest. This method is vulnerable to human bias and can only ever hope to preserve a tiny sliver of the digital universe.
Contrast this with the River approach, epitomized by organizations like the Internet Archive. Here, the strategy is one of automated, bulk capture. Web crawlers, like digital dredges, constantly traverse the internet, scooping up petabytes of data with a goal of comprehensiveness over curation. This is preservation at a scale that tries to match the web’s own boundless growth. The strength of the River is its potential for serendipity and its resistance to selective memory. It preserves the mundane, the commercial, the personal blog, the broken link—the entire digital ecosystem, not just the parts deemed important by a committee. It captures the context in which the "important" sites existed.
But the River has its own profound weaknesses. The resulting archive is a torrent of data, often with inconsistent quality. Captures can be incomplete, missing embedded media or failing to execute complex JavaScript, leaving behind a ghost of a page. More critically, the sheer volume makes the archive nearly impossible to fully index or comprehend. Finding a specific, non-famous piece of information in this deluge can be like trying to find a single, specific fish in the Mississippi. The scale is its salvation and its curse.
The choice isn't about picking a winner. We need both. The Granary gives us depth, context, and vetted quality for specific, high-value historical materials. The River gives us breadth, a bulwark against total loss, and a raw, unvarnished record of our digital culture. The true challenge for our era is figuring out how to let these two models inform each other. How can we build tools that allow scholars to fish meaningfully from the River? And how can the focused efforts of the Granary help us understand what to look for in the torrent? Our digital legacy depends on finding a balance, on building storehouses while also learning to navigate the endless flow.
Notes & further reading
A few pages I came back to while writing this: