The Tyranny of the Open: When Data Availability Obscures Understanding

In the world of open data and digital preservation, our guiding star has long been availability. The rallying cry is familiar: "Set the data free!" We champion initiatives that post massive datasets online, celebrate when public records are digitized, and build elaborate archival systems to ensure that information doesn’t simply vanish into the digital ether. This is presented as an unambiguous good, a victory for transparency and research. But what if this obsessive focus on making data available is, in some cases, making it less, not more, understandable?

We have become so proficient at the mechanics of preservation—crawling, scraping, storing, and serving petabytes of information—that we risk neglecting the more subtle art of contextualization. A terabyte of raw government expenditure data uploaded to a portal is available, yes. But without the accompanying internal memos, the meeting minutes that explain budget anomalies, or the procedural manuals that define the categories, it is also largely inscrutable. We have preserved the skeleton but discarded the nervous system that gave it purpose and meaning. The data is open, but its story is locked away.

This is the tyranny of the open: the creation of a vast, silent library where all the books have blank covers and no one has written an index. We are amassing a digital landscape of orphaned datasets, divorced from the workflows, the arguments, the human decisions, and the institutional quirks that generated them. A future historian might have access to every public Tweet from a certain era, but without the surrounding cultural ephemera, the private reactions, the un-crawled web pages those tweets responded to, the data becomes a cacophony without a score. We are preserving the “what” with astonishing fidelity, while the “why” and the “how” are left to rot.

From Archives to Arguments

The common advice is to prioritize the raw data, under the assumption that context can be added later or that researchers will somehow reconstruct it. This is a dangerous gamble. Context is fragile. It lives in file structures that are flattened during migration, in software interfaces that are never saved, and in the memories of people who move on. By the time a researcher stumbles upon a pristine CSV file from 2025, the living knowledge required to interpret its most crucial columns may already be extinct.

This isn’t an argument against open data; it’s an argument for a more thoughtful, expansive definition of what constitutes the “record.” Preservation shouldn’t end at the dataset’s edge. We need to champion the capture of what archivists call the “function and mandate”—the business rules, the software environments, the email threads debating data quality. This is messier, more complex work than simply uploading a file. It requires us to preserve not just the output, but a slice of the process.

True openness is not just about providing access to a digital artifact. It’s about providing access to understanding. Our goal should be to build archives that are not merely repositories of data, but repositories of arguments, of logic, of the rich and messy human activity that breathes life into ones and zeroes. Otherwise, we risk creating a future that knows everything about our past and understands nothing.

Notes & further reading

A few pages I came back to while writing this: