The Cathedral and the Bazaar of the Web: On the Two Faces of Wayback Machine Captures

I spend a fair bit of time in the Internet Archive’s Wayback Machine, often with the quiet desperation of someone trying to find a lost quotation or a screenshot of a piece of web art that has vanished. Over the years, my view of this incredible archive has shifted. I no longer see it as a single, monolithic entity, but as a collection shaped by two distinct, almost contradictory, methods of capture. One is a method of meticulous planning, the other a spontaneous act of the crowd. One builds a cathedral, the other thrives like a bazaar.

The ‘cathedral’ approach is the work of official, institutional web archiving. Think of national libraries, academic consortia, and other large organizations. They don’t browse; they curate. They compile lists of ‘important’ websites—government portals, major news outlets, accredited university domains—and set their crawlers to methodically harvest them on a strict, predictable schedule. The result is a deep, structured archive. You can visit the home page of a major newspaper on the first of every month for the last decade and find a perfectly preserved snapshot. It is preservation by design, a grand architectural plan for the digital memory of a society. The data is clean, the provenance is clear, and the selection, while necessarily narrow, is defensible. It is a record built for historians of the future.

The ‘bazaar,’ in stark contrast, is the sprawling, chaotic, and wonderfully democratic archive built by us. Every time an individual user, like you or me, types a URL into the Wayback Machine’s search bar and clicks ‘SAVE PAGE NOW,’ we are contributing to the bazaar. This archive is not planned; it is a reaction. It’s the personal blog post saved moments before a server migration, the obscure fan forum page archived by a member who fears a moderator’s purge, the product page saved for a price comparison, the piece of digital activism captured as it unfolds. This record is uneven, idiosyncratic, and deeply human. It captures the web not as an institution sees it, but as individuals experience it.

Each method has its blind spots. The cathedral, for all its order, misses the fringe, the ephemeral, the personal. It archives the official press release but not the ensuing conversation on a now-defunct social media platform. The bazaar, meanwhile, is a patchwork. It’s prone to gaps, to incomplete page saves missing their CSS, to the whims of what a few thousand strangers found compelling on a Tuesday afternoon. It’s a record of attention, not of duty.

What’s fascinating is that these two archives, built for different reasons and by different actors, now coexist within the same domain. They are two distinct lenses on our digital past. The cathedral gives us the official timeline; the bazaar gives us the anecdotes, the marginalia, the lived experience. One tells you what was published; the other often tells you why it mattered enough for someone, somewhere, to click ‘save.’ To understand the true texture of a lost web, you need to consult both. The cathedral provides the grand narrative, but the bazaar holds the soul.

Notes & further reading

A few pages I came back to while writing this: