The Digital Domesday: What If We Could Query the Entire Web of 1998?
Imagine a question about the past, not of kings and treaties, but of dial-up modems and blinking cursor prompts. What was the average file size of a website’s homepage in the autumn of 1998? Which animated GIF was the most ubiquitous? How many personal sites mentioned the release of a specific album or the finale of a popular television show? These aren’t whimsical curiosities; they are data points that could paint a startlingly precise portrait of a cultural moment. The tantalizing, frustrating truth is that while we have the raw material to answer them—petabytes of it in archives like the Internet Archive—we largely lack the means to ask.
We have preserved the individual pages, the ‘books’ of the early web, but we have not yet built the card catalogue for this vast, non-linear library. Web archiving has, out of necessity, focused on the heroic act of capture, of saving the artifacts from the digital deluge. This has given us a sprawling, silent museum where every exhibit is stored in a separate room. To find something, you must already know the room number—the exact URL. Browsing the collection as a whole, asking questions of the entire corpus, remains a monumental technical challenge.
The Chasm Between Storage and Understanding
The core of the problem is one of scale and structure. A web archive is not a single, queryable database. It is a mind-boggling collection of individual WARC files, each containing the raw data of captured web pages, along with their images, stylesheets, and other assets. To ask a question of the whole, one would first need to process all of it: extracting the text, parsing the HTML, identifying the links, and indexing every single element. For a snapshot of the entire web from a single year, this represents a computational task of almost unimaginable proportions, a modern-day equivalent of the original Domesday Book survey of 1086.
Yet, the potential reward is a new form of historiography. Historians of the future wouldn’t just read the diary of one person from 1998; they could analyze the linguistic patterns of a million LiveJournal posts to track the spread of slang. They could map the network of links between early academic websites to see how ideas traveled before social media algorithms dictated our paths. They could study the evolution of design not by looking at a handful of preserved flagship sites, but by quantifying the adoption of table-based layouts versus early CSS experiments across the entire web.
This is the next great frontier for digital preservation: moving beyond saving the bytes to making them speak to each other. It’s the shift from building the archive to building the tools that can listen to its whispers. The 1998 web is a frozen continent, and we have only just begun to chart its coastline. The real discovery awaits when we can finally journey inland and ask it what it knows.
Notes & further reading
A few pages I came back to while writing this:
- Port St Lucie, FL
- The Cartographer of the Unseen: On the Keepers of the UK Web Archive
- Tallahassee, FL
- The Two Temples: Comparing the Internet Archive and Wikipedia as Guardians of the Web
- Tampa, FL
- The Humble Semicolon: A Typographic Fossil in the Data Stream
- Augusta, GA
- Columbus, GA
- Savannah, GA
- Honolulu, HI
- Cedar Rapids, IA
- Des Moines, IA
- Boise, ID