The Digital Paperweight: When Open Data Becomes Too Big to Use

There’s a pervasive belief in our corner of the digital world: that more data is inherently better. The logic seems unassailable. If open data is good, then a torrent of it must be a deluge of goodness. Governments and institutions proudly announce the release of massive datasets—terabytes of sensor readings, decades of transaction logs, comprehensive satellite imagery—as if sheer volume were the ultimate measure of transparency. We celebrate the scale, the ambition, the potential. But in doing so, we risk creating a new kind of obscurity. I call it the digital paperweight: a dataset so colossal and unwieldy that its very size becomes a barrier to understanding, effectively silencing the voices it was meant to empower.

The promise of open data is not just about availability; it’s about usability. It’s about a journalist tracing patterns of civic spending, a community group tracking local air quality, or a student writing a paper. These actors, the lifeblood of a functioning public sphere, often lack the computational firepower and technical expertise to wrestle with data measured in petabytes. They don’t have access to cloud computing clusters or the time to write complex distributed processing scripts. When faced with a dataset that requires a supercomputer just to open, the typical user is not empowered—they are excluded. The data becomes a monument to openness that no one can actually touch, a gesture that looks good in a press release but fails in practice.

This is a failure of curation, not of technology. In our rush to dump everything ‘into the open,’ we’ve skipped the essential, human step of making meaning. An archivist wouldn’t simply bolt a new wing onto a library, fill it with unsorted books in no particular order, and declare the collection accessible. They would catalog, index, and create finding aids. They would produce summaries and guides. Raw data dumps are the equivalent of that un-curated warehouse. They contain immense value, but without thoughtful organization and the creation of entry points, that value remains locked away.

The solution isn’t to stop releasing large datasets. It’s to couple their release with a commitment to what I’d call ‘scaffolded access.’ This means providing not just the raw firehose, but also manageable samples, pre-aggregated summaries, intelligible APIs for common queries, and clear documentation written for humans, not just for machines. It means thinking about the journey of a curious citizen, not just the storage capabilities of a server farm. The goal should be to build a ladder that allows people to climb into the data at a level that matches their resources and questions.

True openness is not measured in terabytes but in the number of people who can actually derive insight. A single, well-documented, and accessible dataset about local park maintenance budgets can do more for civic engagement than a petabyte of un-processed satellite data that only a handful of specialists can interpret. If we are not careful, our grandest open data projects risk becoming the digital age’s answer to the forgotten warehouse at the end of Raiders of the Lost Ark—a vast repository of priceless artifacts, lost to practical use, their stories untold. Let’s build libraries, not just warehouses.

Notes & further reading

A few pages I came back to while writing this: