The Case for the Messy, Unscrubbed Dataset
We spend a lot of time, and rightly so, talking about cleaning data. We champion interoperability, standardised schemas, and pristine, analysis-ready spreadsheets. The common wisdom is clear: data must be wrangled, tamed, and groomed before it can be truly "open" and useful. This is the advice given at every conference and in every primer. But I want to challenge a core, unexamined part of that directive. What if, in our zeal to sanitise, we are performing a kind of archival pre-censorship? What if the most valuable part of a dataset is precisely its glorious, confounding mess?
Consider a municipal database of public works requests. The cleaned version has neat dropdowns: "Pothole," "Streetlight Out," "Graffiti." The raw intake, however, is a text field filled with human urgency. It contains misspellings, local landmark names that never made it to a city map, descriptions of potholes "big enough to lose a bicycle in," alongside complaints about noisy neighbors and lost cats. The cleaned data tells you where to send the asphalt truck. The messy data tells you about the life of the street, the vernacular of a neighborhood, the ancillary concerns that citizens bundle with their civic requests. It is a social record, not just a logistical one.
Our obsession with clean data often assumes a single, known future use-case: statistical analysis, machine learning ingestion, API consumption. But the future is a poor archivist. We cannot know what questions will be asked in fifty years. Historians won’t be mining our perfectly normalized SQL tables for aggregate counts; they’ll be hunting for the anomalies, the outliers, the human fingerprints in the metadata. They’ll want the draft versions, the abandoned fields, the comments left by a long-gone intern explaining why a certain record looks "funny." This context—the lint in the pocket of the data—is the first thing scrubbed away in the name of cleanliness.
This isn't an argument against cleaning data for operational use. By all means, build your efficient systems. But when we archive, when we preserve and proclaim something as an open public record, we must make a conscious choice. Are we archiving a refined product, or are we archiving a source? The push for always-clean open data privileges the former, treating data as an end-point commodity. It risks creating a future digital history composed solely of press-ready summaries, devoid of the friction that signals genuine human and institutional activity.
Preserving the Process
The counterintuitive proposal, then, is this: when releasing or preserving data, we should prioritize the "raw feed" alongside, or even instead of, the cleaned version. Document the cleaning process meticulously, yes, but keep the original, warts-and-all capture. Call it the "archival original," and treat it with the same respect a museum gives to an artist's preparatory sketches. The value is in the lineage, in the ability to see not just the conclusion, but the path taken to get there. In web archiving, we keep the WARC with all its HTTP headers and failed requests—not just the pretty, re-rendered page. Why should our structured data be any different?
In the end, a perfectly clean dataset is a closed argument. A messy one is an open conversation. It admits fallibility, captures nuance, and invites reinterpretation. For those of us committed to true openness, our goal shouldn't be to present a spotless, finished facade. It should be to preserve the valuable, noisy, and wonderfully inconvenient truth of how things actually were.
Notes & further reading
A few pages I came back to while writing this:
- Columbus, GA
- The Granite Ledger: How a 19th-Century Statistician Built an Open Data Ark
- Savannah, GA
- The Scent of Celluloid: On Finding a Lost Year in a Databank
- Honolulu, HI
- The Hum of the Machine: On the Permanence of the Temporary
- Cedar Rapids, IA
- Des Moines, IA
- Boise, ID
- Aurora, IL
- Chicago, IL
- Joliet, IL
- Rockford, IL