The Fiction of Findability: On the Assumption That Public Data is Effectively Searchable

There is a pervasive belief in the world of open data and public records that publication equals accessibility. We imagine that once a dataset is uploaded or a document is scanned and placed on a government server, its journey is complete. The information is now 'out there,' democratically available to any citizen with an internet connection. This belief, while comforting, obscures a more stubborn reality. The true bottleneck in our modern information ecosystem is not access, but findability. We have thrown open the doors to vast libraries, only to find that there are no card catalogs, and the books are arranged in an order known only to the original shelvers.

This is the fiction of findability. It’s the unexamined assumption that data, once made public, can be effectively discovered and used by those who need it. In truth, many digital repositories are structured for the convenience of the hosting institution, not the searching public. Records are often trapped in labyrinthine departmental websites, labeled with obscure project codes or internal jargon that is meaningless to an outsider. A researcher looking for environmental impact studies might need to know the precise name of the permitting division that approved a project twenty years ago—a piece of procedural knowledge that is itself a form of hidden data.

The problem is compounded by the poverty of metadata. A thousand-page PDF of meeting minutes might be posted with a filename like 'Council_Minutes_2023_10_v2_Final.pdf.' While technically a description, this tells us little about the actual content. What key decisions were made? Which ordinances were debated? Without descriptive tags, summaries, or structured data, each document becomes a digital monolith that must be manually scaled and explored. This is a task for which algorithms are poor substitutes and human effort is prohibitively time-consuming. The data is open, but it is not legible.

Furthermore, we have an over-reliance on primitive search functions that are easily defeated by the very formats we use. The search box on a typical public records portal is often a blunt instrument, incapable of parsing the contents of scanned images or complex spreadsheet tables. It treats a keyword as a string of characters, not a concept. A search for 'green space allocation' will miss documents that discuss 'parkland dedication' or 'urban canopy cover,' simply because the specific term of art is different. The semantics of the data—its actual meaning—are lost in translation between the bureaucrat who created it and the citizen trying to find it.

This fiction has consequences. It creates a two-tiered system of information access. The savvy insider—the journalist, the policy analyst, the seasoned activist—learns the tricks. They know which acronyms to use, which departments hold which records, and how to construct a query that the primitive search box will understand. The average citizen, or the small community group with limited resources, is left with the illusion of access. They can see the library exists, but they cannot find the book they need. True openness, then, is not just about making data public. It is about the painstaking, unglamorous work of making it genuinely findable: crafting meaningful descriptions, investing in semantic search, and building interfaces that bridge the gap between public need and bureaucratic language. Until we confront this fiction, our digital repositories will remain only half-open.

Notes & further reading

A few pages I came back to while writing this: