The Tyranny of the PDF: How a Standard Became a Barrier to Open Data
We are told we live in an age of open data. Governments, institutions, and corporations proudly proclaim their commitment to transparency by releasing troves of public records. But too often, this commitment is betrayed by the very format chosen for release: the PDF. What began as a portable, reliable way to share formatted documents has become, in the context of public data, a digital cage.
The received wisdom is that a PDF is a safe, universal standard. It looks the same on every screen, prints predictably, and preserves the intended layout. For a final report meant for human eyes only, this is perfectly adequate. But when a PDF is used to publish data that should be queried, analyzed, or cross-referenced—budget figures, spending records, legislative votes—it ceases to be a standard and becomes a barrier. It is a format of presentation, not of participation.
This creates a two-tiered system of access. A journalist, researcher, or curious citizen seeking to analyze a city’s procurement data might find it in a 300-page PDF council minutes packet. The information is technically public, yet utterly imprisoned. To free it, one must perform the digital equivalent of manual labor: copying by hand, or relying on often-faulty optical character recognition (OCR) software that can mangle numbers and dates. This process introduces error, demands significant time, and effectively taxes the public’s right to know.
The Illusion of Accessibility
This reliance on the PDF creates an illusion of accessibility. The checkmark for "public record" is satisfied, but the spirit of open data—machine-readable, easily reusable information—is utterly defeated. It is a form of transparency theater. The data is visible, like fish behind glass, but entirely out of reach.
The argument for PDFs often hinges on convenience for the publisher, not usability for the public. It is far easier to dump a scanned printout or a Word document converted to PDF into a portal than it is to structure data properly into a CSV or JSON file. This convenience comes at a profound cost to civic utility. It privileges the aesthetics of officialdom over the messy, vibrant utility of raw data that can be sorted, filtered, and remixed.
True open data requires a mindset shift from publishing documents to publishing datasets. It asks publishers to consider not just whether information can be seen, but what can be done with it. The humble CSV file may lack the official seal and fancy fonts of a PDF, but it embodies a far deeper commitment to openness. It is an invitation to engage, not just to observe. Breaking free from the tyranny of the PDF means demanding data we can actually use, not just data we can technically see.
Notes & further reading
A few pages I came back to while writing this:
- Irving, TX
- The Quiet Art of the Web Archive Capture: A Field Guide to Snagging Elusive Interactive Maps
- Killeen, TX
- In Praise of the Broken Link: Why Digital Ruins Matter
- Laredo, TX
- The Broken Spoke: Reconstructing a Lost Network from Digital Ghosts
- Lubbock, TX
- Mcallen, TX
- Mckinney, TX
- Mesquite, TX
- Midland, TX
- Pasadena, TX
- Plano, TX