The Broken Ladder: On the Illusion of Perfect Open Data Formats
If you've spent any time in the world of open data, you've heard the mantra: choose an open format. It’s the foundational advice, the first rung on the ladder to a transparent and durable digital public square. We are told to champion CSV over XLSX, PDF/A over a scanned image, plain text over a proprietary database. The logic is unimpeachable: open formats are machine-readable, future-proof, and vendor-neutral. They are, we assure ourselves, the responsible choice. But what if this well-intentioned dogma has become a convenient excuse for inaction, a ladder leading to a platform of pristine, empty perfection where no real data ever lives?
We have become so obsessed with the ideal container that we often forget the messy, vital importance of the contents. Consider a municipal government that releases its annual budget. The open data advocate insists it must be a tidy CSV file, with perfectly normalized columns and standardized codes. The town clerk, overwhelmed and under-resourced, finds the process of converting their internal spreadsheet too complex, too time-consuming. The result? The data isn’t released at all. The perfect has become the enemy of the good, or even the tolerable. In our quest for archival purity, we’ve built a gate that keeps more data locked away than it sets free.
This fixation on format is a form of premature preservation. It presumes a future where our primary challenge will be technical interoperability, rather than simple, brute-force access. But the more immediate threat to public records isn't obsolescence; it's obscurity. A budget trapped in a proprietary spreadsheet from 2010 is, without question, a problem. But it is a solvable problem. A budget that was never released because it couldn't be perfectly formatted is a total loss. It is a null set in the public ledger, a silence we have politely engineered for ourselves.
The Tyranny of the Pristine File
This creates a perverse incentive. For the data publisher, the pressure to offer a 'perfect' open dataset can be paralyzing. It's safer to do nothing than to risk the criticism of data activists pointing out a flawed schema or a non-standard date field. Meanwhile, the much messier, but far more abundant, data flows all around us are ignored. The thousands of PDF meeting minutes, the image-based press releases, the sprawling, unstructured council agendas—these are treated as second-class records, not worthy of archiving efforts until they can be 'properly' converted. But this is where the actual history of governance is happening, right now, in real-time.
What we need is a shift in priority, from format fetishism to capture pragmatism. The primary goal should be to get the information into the public domain, in whatever form it currently exists. The work of standardization and refinement should come after capture, not before. A PDF, even an image-based one, can be OCR'd. A strange XML feed can be parsed. A convoluted spreadsheet can be untangled. These are tractable problems for archivists and programmers. The one problem we cannot solve is retrieving a record that was never preserved at all because its format was deemed unworthy.
Let's stop treating open formats as the starting gate and start treating them as a finish line. The first, most critical act is to build the digital dustpan wide enough to catch everything. We can sort and clean the fragments later. A flawed, messy, imperfect public record is infinitely more valuable than the flawless emptiness of a perfectly formatted, hypothetical one. The ladder of open data is broken if the first rung is set too high for anyone to reach. It's time we started climbing the pile of usable rubble instead.
Notes & further reading
A few pages I came back to while writing this: