The Digital Dustpan: A Simple Guide to Scraping Public Meeting Transcripts
We often talk about web archiving in grand, technical terms—crawling entire domains, preserving complex interactive media. But some of the most vital public records are hiding in plain sight, published in a consistently dull format that makes them perfect for a simple, targeted rescue mission. I’m talking about the text transcripts of public meetings: city councils, school boards, planning commissions. These are the raw, unvarnished minutes of civic life, and they are often published on fragile, poorly designed government portals where a single website redesign can erase years of discourse.
This isn’t about building a complex scraper. It’s about using a humble, often-overlooked tool: the command-line utility `wget`. Think of it not as a bulldozer, but as a precise digital dustpan. Your goal isn’t to download the entire city’s website, but to carefully sweep up a specific set of files before they get lost.
The technique relies on a beautiful bit of consistency. Most municipalities post these transcripts as PDFs, and their URLs often follow a predictable pattern. For example: `https://www.cityname.gov/meetings/2024-01-15_minutes.pdf`, `https://www.cityname.gov/meetings/2024-02-19_minutes.pdf`. The date changes, but the structure remains. This predictability is your key.
Here’s the simple command, which you can run in your terminal. First, navigate to a directory where you want to store these public records. Then, craft your `wget` command like a incantation:
wget --recursive --level=1 --accept pdf --no-parent --wait=2 https://www.cityname.gov/meetings/
Let’s break down this spell. `--recursive` tells it to follow links. `--level=1` ensures it only goes one layer deep, avoiding a crawl of the entire city website. `--accept pdf` is the crucial filter; it will only download files ending in .pdf. `--no-parent` keeps it from wandering up the directory tree. And `--wait=2` is an act of good citizenship, adding a two-second pause between requests to avoid hammering the public server.
This one command will systematically work through the meeting directory page, find all the linked PDFs, and download them to your machine. You’ve just created a local, stable archive of civic dialogue. The files are now safe from link rot, site migrations, or accidental deletion. You can run this same command every month to harvest the latest minutes, building a continuous record.
The beauty of this method is its simplicity and its focus. It doesn’t require a programming language or a fancy framework. It uses a decades-old tool to perform a crucial act of preservation. It’s a small, deliberate gesture against entropy, ensuring that the words spoken in a public forum on a Tuesday evening remain a readable, accessible record for anyone who cares to look. It’s not flashy, but it is profoundly important.
Notes & further reading
A few pages I came back to while writing this: