The Myth of the Immortal PDF: Why Our Default Format Fails the Future
We have been sold a lie, a comfortable and convenient one. For decades, the PDF has reigned as the undisputed champion of digital document preservation. Governments, institutions, and corporations worldwide treat it as the final, unassailable format for public records. Click ‘Save As PDF’ and, the thinking goes, you have etched your information in digital stone, preserving it perfectly for eternity. This is the received wisdom we must now challenge. The PDF is not a preservation format; it is a presentation format, and our blind faith in its permanence is creating a silent crisis of inaccessibility.
The core of the problem lies in the very thing we celebrate about PDFs: their fixed nature. A PDF is designed to look the same on every screen, to replicate the experience of a printed page. This rigidity is its greatest weakness for long-term archival purposes. The information within a typical PDF is often a black box of unstructured data—a flattened image of text, not the text itself. For a human reader, this might be sufficient, but for the systems and algorithms that will need to parse, index, and understand these records in the future, it is nearly useless. It is a photograph of knowledge, not knowledge itself.
Consider a simple public record, like a municipal budget. Saved as a PDF, the numbers within it are locked away. You cannot easily export the data to a spreadsheet for analysis. You cannot have a tool automatically compare line items across years. The data is trapped, requiring manual—and error-prone—re-entry to be truly utilized. This fundamentally violates a key principle of open data: machine-readability. We are creating vast archives of documents that are human-readable only through considerable effort, future-proofing nothing but the document’s visual layout, often at the expense of its actual content.
The issue compounds over time. Proprietary formats evolve, and while PDF is a standard, it is not immune to obsolescence. How many of us have encountered a PDF from two decades ago that renders incorrectly or not at all in a modern viewer? More critically, as assistive technologies for the visually impaired become more advanced, a static PDF can be a significant barrier, its accessibility wholly dependent on how well it was constructed at the moment of its creation, a moment now frozen in the past.
This is not an argument to abandon the PDF, but to dethrone it as our default. The true path to preservation is layered. The archival version of a public record should be its raw, structured data—a CSV file, an XML feed, plain text. The PDF can exist alongside it as a convenient, human-friendly snapshot. By mistaking the snapshot for the subject, we are building a future archive that is beautiful to look at but impossible to truly read. We are preserving the container and letting the content wither inside.
Notes & further reading
A few pages I came back to while writing this:
- Pasadena, CA
- The Unseen Web: How to Archive a Single, Crucial Social Media Post
- New Haven, CT
- The Poisoned Well: When Open Data Corrupts the Record
- Stamford, CT
- The Cuneiform Cache: Data Carriers of the First Empire
- Washington, DC
- one area's overview
- a practical rundown
- Little Rock, AR
- Gilbert, AZ
- Peoria, AZ
- Surprise, AZ