The Poisoned Well: When Open Data Corrupts the Record
We are living in the era of the open data mandate. From city budgets to scientific research, the prevailing wisdom is clear: release the data, as-is and in bulk, to the public. This is celebrated as an act of radical transparency, a way to empower citizens and researchers alike. But what if this well-intentioned act is sometimes poisoning the well of public understanding itself?
The common advice is to prioritize the raw data above all else. Processed reports, summarized findings, curated dashboards—these are often viewed with suspicion, as potential vehicles for spin or editorial bias. The raw CSV file, by contrast, is seen as pure, unadulterated truth. This belief is so entrenched that we rarely pause to consider a dangerous side effect: the wholesale release of raw data can actively corrupt the historical record by stripping it of its native context and credibility.
Consider a municipal government releasing decades of public works tickets—a raw dump of service requests, from potholes to broken streetlights. The dataset is a treasure trove. But what it likely lacks are the critical footnotes that gave each entry its meaning. It won’t record that the spike in pothole complaints in Ward 3 during April 2015 wasn't due to failing infrastructure, but to a proactive, well-publicized campaign by a new councilor to encourage reporting. A future data analyst, seeing only the numbers, might confidently conclude that Ward 3’s roads abruptly deteriorated that spring, perpetuating a statistical fiction.
This is the counterintuitive corruption. The data, in its rawest form, lies by omission. The very act of "opening" it by stripping away the bureaucratic scaffolding—the memos, the policy changes, the internal communications that explain anomalies—creates a more accessible but fundamentally less truthful artifact. We are preserving the seed but discarding the instructions for its growth, ensuring that anyone who plants it will cultivate a distorted version of reality.
The Burden of the Unguided Dataset
This problem is compounded by the sheer effort required to re-contextualize this data. The promise of open data is that the crowd—journalists, academics, civic hackers—will do this work. But this imposes a massive and often impossible burden. How is an independent researcher to know that a particular code was retired in 2010, or that two departments merged their record-keeping systems in a way that created a phantom drop in certain types of requests? The original custodians of the data possess this institutional memory; by releasing the data without it, they are outsourcing a forensic investigation with no starting clues.
In our rush to open the vaults, we have conflated accessibility with accountability. True transparency isn’t just about providing the numbers; it’s about providing the narrative that makes those numbers honest. A PDF summary report from 2008, which we might dismiss as "closed" data, could contain the crucial paragraph explaining a methodological change that renders the raw data from before and after that date incomparable. By discarding these "curated" artifacts in favor of pure bulk data, we are, paradoxically, choosing a less complete and more easily misconstrued record.
The solution is not to stop releasing data, but to challenge the dogma of "raw above all." Preservation and openness must include what we might call "contextual metadata"—the memos, the press releases, the meeting minutes that explain the data’s creation and evolution. It means preserving and linking to the summarized reports alongside the spreadsheets. It acknowledges that data is not born in a vacuum; it is a product of a specific time, place, and human process. To archive the data without its story is to preserve a body without a soul, a relic that future generations may study with immense precision only to arrive at a profoundly incorrect conclusion.
Notes & further reading
A few pages I came back to while writing this:
- Stamford, CT
- The Cuneiform Cache: Data Carriers of the First Empire
- Washington, DC
- The Ghost in the Spreadsheet: A Memory of My Grandfather's Last Public Record
- one area's overview
- The Unwritten Ledger: On the Silent Loss of Unrecorded Hyperlinks
- a practical rundown
- Little Rock, AR
- Gilbert, AZ
- Peoria, AZ
- Surprise, AZ
- Elk Grove, CA
- Pasadena, CA