The Tyranny of the PDF: Why 'Accessible' Formats Can Obscure Public Data

In the world of public records and open data, the PDF is often hailed as a victory. A city council publishes its minutes, a federal agency releases a report, a research institute shares its findings—all as neat, tidy PDFs. The logic seems unassailable: it’s a standard format, it preserves layout, and anyone can open it. It’s the go-to solution for making documents ‘accessible’ to the public. But this received wisdom, this default setting of digital civics, harbors a quiet tyranny. It confuses publication with true accessibility, and in doing so, it builds a new kind of wall around public information.

The Illusion of Openness

The PDF is, at its heart, a presentation layer. It is designed to look exactly the same on every screen, a digital snapshot of a printed page. This is its great strength for fidelity, and its fatal flaw for data. When a 300-page environmental impact report is released as a single, image-heavy PDF, it is effectively a black box. The data within—the water quality tables, the species counts, the emissions projections—are locked in a visual prison. A human can read it, painstakingly. But a researcher cannot efficiently analyze it. A journalist cannot quickly sort it. A civic technologist cannot remix it to create a public dashboard. The data is present, but it is not operable. It is published, but not truly open.

This creates a two-tiered system of access. Those with resources—law firms, large nonprofits, corporate interests—can afford the time or the OCR software and data-entry labor to crack these documents open. The average citizen, the small watchdog group, or the curious academic is left with the daunting prospect of manually transcribing tables or wrestling with clunky ‘export to text’ functions that mangle the structure. The PDF, in its passivity, becomes a tool that maintains the status quo of information asymmetry, all while wearing the benign mask of public service.

The deeper issue is one of intent. Choosing the PDF as the final, definitive format for public records often signals a completion of duty: “We have made it available.” The work stops at the point of publication, not at the point of public utility. It prioritizes the convenience and workflow of the publishing institution (which often works with document editors geared toward print) over the needs of the data-consuming public. It’s a digital extension of the old model of placing a physical binder on a shelf in a records office during business hours. You can see it’s there, but actually using it requires disproportionate effort.

True open data isn’t just about putting files on a server. It’s about enabling understanding, analysis, and reuse. This means complementing the presentation PDF (if it must exist) with the underlying structured data: the budgets as CSV files, the permits as JSON feeds, the legislation as machine-readable XML. This isn’t a futuristic ideal; it’s a matter of shifting the default. The tyranny of the PDF isn’t in its existence, but in its dominance as the supposed end-point of transparency. To liberate public data, we must challenge the wisdom that a document made for human eyes is the same as data made for public use. The next frontier of civic access isn’t about getting the document, but being able to truly converse with the information inside it.

Notes & further reading

A few pages I came back to while writing this: