The Myth of the Universal Format: On PDFs and the Illusion of Preservation
We have been sold a bill of goods, and it’s wrapped in the comforting, familiar icon of a document with a folded corner. The PDF, or Portable Document Format, has become the default standard for publishing everything from academic papers to government reports, from utility bills to public meeting minutes. Its promise is one of perfect preservation: a file that looks the same on any screen, anytime, forever. This promise, however, is a siren song leading our public records into a silent, unsearchable trap.
The received wisdom is that by saving something as a PDF, we are archiving it. The data is safe. The record is kept. But this confuses visual fidelity with functional preservation. A PDF is often little more than a digital photograph of a document. It prioritizes the *look* of information—the font, the layout, the pagination—over the information itself. The text within a scanned PDF is, to a machine, just a collection of pixels. It is an image, not data. It cannot be easily searched, indexed, or computationally analyzed without first undergoing the often flawed process of Optical Character Recognition (OCR).
This creates a profound accessibility problem. A thousand-page budget report released as a single, massive PDF of scanned pages is not an open record; it is a locked vault. Researchers, journalists, and curious citizens are forced to manually sift through it, rather than using tools to quickly find, aggregate, and cross-reference the data within. It places a burden of labor on the information seeker that is antithetical to the very spirit of open data.
The issue deepens with the rise of more ‘modern’ PDFs built from word processors. While these contain actual text, they are often structured not for clarity but for visual layout. Sentences are broken into disparate text boxes, tables are rendered as artistic elements rather than structured data, and critical context is lost. The information is there, but it is not semantically organized. It’s a heap of words, not a structured document.
Our blind faith in the PDF as an archival format represents a failure of imagination and a prioritization of convenience over utility. We have chosen the format that is easiest for the *publisher* to create—the ‘Print to PDF’ button—rather than the format that is most useful for the *public* to consume. True digital preservation and open data require machine-readable, structured formats like CSV, JSON, or XML for data, and properly tagged HTML for documents. These formats separate content from presentation, ensuring the information remains accessible, malleable, and alive for future tools and inquiries.
By accepting the PDF as the final word, we are not preserving our records for the future; we are merely taking a very high-quality picture of them as they fade away. The true public record shouldn’t be a snapshot. It should be a seed.
Notes & further reading
A few pages I came back to while writing this:
- a useful directory
- Hunting the Deleted: A Beginner's Guide to Using Change Detection for Web Archiving
- one area's overview
- The Flaw in Perfect Memory: A Case for the Irretrievable
- a helpful reference
- The Scrapbook Queen and the Unread Ledger: On Memory, Purpose, and the Public Record
- a place-by-place guide
- a practical rundown
- a local resource
- a regional guide
- a nearby resource
- Anchorage, AK
- Birmingham, AL