The Scribe and the Cartographer: Two Paths to Preserving a Website
What does it mean to keep a website? The impulse is simple enough: we feel a page is important, or beautiful, or representative of a moment, and we want to ensure it doesn’t vanish into the 404 void. But the methods we choose to answer this call reveal profoundly different philosophies about what, exactly, we are trying to save. Two dominant approaches—the Wayback Machine and the Wget command—embody this division. One aspires to be a universal scribe, the other a precise cartographer.
The Internet Archive’s Wayback Machine is our most public-facing scribe. Its mission is grand: to take a snapshot of the entire, sprawling web, as often as possible. When you enter a URL into its search bar, you are summoning a chronicler who has been dutifully, automatically transcribing the internet for decades. The result is a timeline, a palimpsest of a site's public life. You see it as a visitor might have seen it, with functional links (to other archived pages), loaded images, and the general feel of the original. The scribe’s goal is fidelity to the experience of the web.
But this approach has its quirks. The archive is not instantaneous; the scribe may have visited on a Tuesday when you wish it had come on a Wednesday. Some complex, interactive, or script-heavy pages are captured imperfectly, like a transcription of a play that notes where the laughter occurred but can’t reproduce the actor’s tone. It’s a miraculous, crowd-sourced memory, but it’s a memory seen from the outside, dependent on the schedule and capability of the archiving bot.
The Cartographer's Precise Map
Contrast this with the approach of a tool like Wget, the quiet workhorse of command-line enthusiasts. Wget is a cartographer. You give it a URL and a set of instructions, and it methodically downloads every file it can find, mapping the topography of the site onto your local hard drive. It doesn’t interpret the experience; it collects the raw materials—the HTML, the CSS, the images, the PDFs. The result is not a timeline to browse, but a static data set, a perfect replica of the site’s structure at a single, precise moment.
The cartographer’s map is complete and self-contained. It doesn’t rely on the Internet Archive’s servers or link out to other archived pages. It is your own private, frozen copy. This is the tool for the researcher who needs to perform textual analysis on a site’s content long after it’s gone, or for the individual who wants to preserve a personal project with absolute certainty. The trade-off is that this map is inert. The dynamic nature of the original is lost, reduced to its constituent parts.
So, which is the truer preservation? The answer depends on what you value. The scribe (the Wayback Machine) preserves context and interconnection, giving us a sense of the web as a living ecosystem, even if some details are blurry. The cartographer (Wget) preserves content and structure with forensic precision, creating a definitive artifact for future study, even if it sacrifices the living breath of the thing.
Perhaps the most robust strategy is to recognize that we need both. We need the grand, collaborative project of the scribe to maintain our collective digital commons. And we need the focused, deliberate work of the cartographer for the specific corners of the web we hold most dear. One writes the history; the other secures the evidence. Together, they offer a more complete answer to the haunting question of what we leave behind.
Notes & further reading
A few pages I came back to while writing this: