The Unseen Hand of the Precise Paraphrase: How Open Data Gets Warped on Its Way to You

We celebrate the open data movement for its promise of unvarnished truth. Raw datasets, direct from the source, are meant to cut through the noise of interpretation and spin. The ideal is a pristine pipe, connecting the well of public information directly to our taps. But what if, on its journey to us, the water is subtly filtered, its mineral content adjusted by an unseen hand? This isn’t a story of malicious data manipulation, but of something more insidious and commonplace: the quiet art of the precise paraphrase.

Consider a municipal database detailing code enforcement complaints. The original, hand-written report from a city inspector might read: "Resident reported a strong, foul odor emanating from neighbor’s property, possibly from accumulated refuse. Observed several overflowing trash bags near the driveway." This is rich with human context. But to fit neatly into a standardized, machine-readable ‘Open Data’ field labeled ‘Complaint Type,’ this narrative must be categorized. A clerk performs the translation. The complex, sensory complaint becomes a sterile, searchable code: TRASH_VIOLATION.

This act of translation is the first and most critical point of warping. The specific anxiety of the resident, the inspector’s own observation—these nuances are shed for the sake of tidy data. The record is now computationally useful, but humanly poorer. The foul odor, the possible source, the relational tension between neighbors; all are lost to the algorithmic ledger. The data is open, but it has already been interpreted for you.

The Cascade of Simplification

The problem doesn't end with the initial data entry. This pre-digested information then flows to journalists, researchers, and civic app developers. A data journalist, on a deadline, queries the database for all ‘TRASH_VIOLATION’ complaints in a specific ward. They write a story about the prevalence of sanitation issues. The story is factually correct, but it misses the underlying narrative of neighborly disputes and specific public health concerns that the original reports contained. The journalist is working with a shadow of the original record, a caricature drawn by the constraints of a dropdown menu.

This cascade of simplification continues. A developer uses the same API to build a ‘See Click Fix’ style app, where citizens can report issues. The app, designed for efficiency, presents users with the same limited set of categories. The cycle reinforces itself. The system, built for openness, now dictates the very language by which we are allowed to describe our civic reality. We begin to see our world through the categories the database understands.

This isn’t an argument against categorization or open data portals. They are powerful tools for transparency. It is, however, a plea for a more sophisticated form of archival humility. True openness must mean more than just providing access to sanitized endpoints. It requires a commitment to preserving the original, messy, narrative-rich records alongside their tidy, datafied twins. It asks us to be not just consumers of the paraphrase, but seekers of the original text. The next time you download a pristine CSV file from a government portal, ask yourself: what was the original language? What human story was flattened to fit that column? The deepest truths in public records are often hidden not in the data itself, but in the gaps between the code.

Notes & further reading

A few pages I came back to while writing this: