The Citizen Archivist's First Trowell: A Practical Guide to Wget
We talk a great deal about the fragility of the web, about link rot and digital decay, but the conversation can often feel abstract. It’s a problem happening somewhere out there, in the vastness of the network, to data we don’t own. But what if you want to get your hands dirty? What if you want to become a custodian, even in a small way, of a corner of the web you care about? The leap from concern to action requires a tool, and one of the most powerful, fundamental tools for this work is also one of the oldest: a command-line program called Wget.
Think of Wget not as a magical preservation wand, but as a archaeologist’s trowel. It’s a simple, precise instrument for carefully excavating digital content. For our purposes, it allows you to download the contents of a website—its HTML, images, and stylesheets—and recreate its structure locally on your own computer. This creates a static snapshot, a frozen moment in time that you can keep safe.
An Unlikely Solution from a Simpler Time
Born in 1996, Wget hails from an era before single-page applications and complex javascript frameworks. This is its greatest strength for the beginner archivist. It works by mimicking the behavior of a web browser, but instead of displaying pages, it saves them. It follows links within a defined scope, methodically building a mirror of the site. While it struggles with highly dynamic, interactive modern web apps, it is exceptionally good at preserving the vast expanses of the web that are still built on humble, link-based HTML—precisely the kind of personal blogs, community forums, and informational sites most at risk of vanishing.
To start, you’ll need access to a command line. On macOS and Linux, it’s built right into the Terminal. Windows users can easily install it or use the Wget that comes with the Windows Subsystem for Linux. The basic command is startlingly simple. To save a single page, you would type: wget https://example.com/my-favorite-blog-post.html. But the real power comes with a few additional flags that instruct Wget on how to behave like a conscientious archivist.
The most crucial flag for our purpose is --page-requisites or -p. This tells Wget to download all the files necessary to display the page correctly—the images, the CSS that gives it style. Without this, you’d just have a bare-bones HTML file. Next, we want to follow links, but not too far. The flag --limit-rate=500k ensures you don’t overwhelm the host server with requests, acting as a polite digital guest. And --wait=2 will pause for two seconds between requests, another gesture of good etiquette.
So, a robust command to create a local archive of a small personal website might look like this:
wget --page-requisites --limit-rate=500k --wait=2 --no-parent https://exampleblog.com/about/
The --no-parent flag is important here; it tells Wget to stay within the /about/ directory and not go crawling up to the root of the entire site, keeping your project focused. When you run this, Wget will work quietly, printing its progress to the screen. Once finished, you’ll have a new folder containing the complete ‘about’ page, ready to be opened in your browser, even when you’re offline. You’ve just built your first tiny archive.
This isn’t a perfect solution for every website, but it’s a start. It’s a direct, hands-on way to engage with digital preservation. That folder on your hard drive is more than a backup; it’s an act of defiance against the web’s inherent transience. It’s you, with a simple tool, saying: this mattered to me, and I will keep it.
Notes & further reading
A few pages I came back to while writing this: