The Illusion of the Readable Record: When Clarity Obscures Meaning

In the worlds of open data and public records, there is a sacred mantra: make it readable. We are told to convert scanned PDFs into machine-readable text, to clean and structure datasets, to translate bureaucratic jargon into plain language. The goal is always to remove friction, to polish the artifact until its information gleams with immediate accessibility. This seems an unassailable good. But what if, in our relentless pursuit of clarity, we are inadvertently sanding off the very texture that gives a record its true context and meaning?

Consider the humble government memo, scanned and then run through an OCR engine. The standard practice is to correct the errors, to fix the misreads of a faded typewriter ribbon or a coffee-stained corner. We produce a pristine text file. Yet, in doing so, we erase the evidence of its physical life. The original scan’s quirks—the smudged letter, the stray pen mark from a recipient, the faint letterhead of a defunct department—were not noise. They were data. They spoke of the document’s journey, its use, its material existence in a world of paper and people. Our 'clean' version presents a fiction: a record born digitally perfect, devoid of history.

The Seduction of the Structured Field

This obsession extends to data. We take messy, sprawling archives and force them into neat, relational databases. We define fields, establish controlled vocabularies, and separate 'data' from 'metadata.' The result is searchable, sortable, and wonderfully efficient. But the original organization—or lack thereof—often held its own intelligence. The order of pages in a clerk's bound ledger, the idiosyncratic filing system of a local office, the handwritten annotations in the margins of a spreadsheet: these are the scars and fingerprints of human process. By imposing our own logical structure, we risk overwriting the native logic of the record’s creators, losing the story of how the information was actually used and understood in its own time.

The push for 'readable' public records often means translating the specialized language of an institution into something a general audience can quickly grasp. But in simplifying the legalese or the technical terminology, we can drain the precise, deliberate meaning the authors intended. A term like “shall” carries a specific, binding weight in a legal context that “must” or “will” does not fully capture. The formal, repetitive phrasing of a regulation isn’t just bureaucratic bloat; it is a painstaking effort to cover contingencies and avoid ambiguity. Our readable paraphrase is, by necessity, an interpretation, and a potentially reductive one.

None of this is an argument for abandoning accessibility or for leaving data in unusable forms. It is, however, a plea for humility and layered preservation. The goal should not be to replace the complex, 'unreadable' original with a clean facsimile, but to preserve both, and to make their relationship clear. The OCR text should live alongside the scan; the structured database should link back to images of the original ledger pages; the plain-language summary should sit beneath the full, unaltered text.

True transparency isn't just about providing the answer. It's about providing the artifact, in all its confounding, messy, human glory, and trusting the public to engage with its full depth. Sometimes, meaning is found not in the clear pane of glass, but in the unique distortions of the original, wavy pane.

Notes & further reading

A few pages I came back to while writing this: