The Unchecked Ledger: On the Assumption of Neutrality in Public Record Sets

There’s a quiet, persistent faith that underpins much of the open data movement: the belief that a dataset, once liberated from its silo and published in a raw, machine-readable format, becomes a neutral artifact. It is data, stripped of opinion, ready for objective analysis. This assumption is our received wisdom, and it is dangerously incomplete. We celebrate the opening of the vault but seldom audit the ledger’s original entries.

Consider a municipal dataset of business licenses. It lists names, addresses, dates, and categories. To the civic hacker, it’s a treasure trove for mapping economic corridors. But who defined those categories? Was a corner store run by a new immigrant consistently classified under the same “retail” code as a decades-old family deli, or did inconsistent clerical judgment create shadow categories? The data is clean, but the process that created it was human, messy, and subject to the biases and bureaucratic whims of its time. The dataset doesn’t record the applications that were discouraged at the counter, or the businesses that operated for years in a grey area never captured by a form. What we have is not a perfect record of commerce, but a perfect record of what the bureaucracy successfully documented.

The Archive of the Permitted

This extends far beyond commerce. Police incident reports, published as open data, are often treated as a straightforward map of crime. Yet they are, first and foremost, a map of reported crime, of policing priorities, of community trust (or lack thereof) in authority. A blank spot on the map is not an area of perfect safety; it is an area of disconnect, of fear of reprisal, or of resigned acceptance. The dataset is silent on its own silences. To analyze it as a pure, neutral truth is to accidentally endorse its inherent perspective—to mistake the archive of the permitted for the complete story.

Digital preservation and web archiving face a parallel pitfall. We strive to save the *what*—the HTML, the images, the styles—with incredible fidelity. But we often fail to capture the *why* of a page’s existence, or the *how* of its journey to being archived. Was this site saved because it was linked by a major institution, or because a lone archivist found its niche community valuable? The corpus of the archived web is not a democratic sample of human expression; it is a collection of judgments about what future generations might find significant. It is a cultural argument masquerading as a library.

This isn’t an argument against open data or web archiving. It is a critique of the innocence we assign to them. A public record set is not a pristine well of truth. It is a fossil. And like any fossil, it tells a story shaped as much by the conditions of its preservation and the anatomy of the creature that made it as by the ecosystem it once inhabited. The work, then, is not done when the CSV file is posted. The real work begins with a critical reading of the ledger itself—questioning its categories, hunting for its absences, and understanding the office, the rules, and the people who first decided what was worth recording, and what was not. The data is open, but our eyes must be open wider.

Notes & further reading

A few pages I came back to while writing this: