The Scent of a Database: What Happens When 'Unstructured' Isn't a Dirty Word

Most discussions about open data and public records revolve around neat, orderly spreadsheets and meticulously tagged documents. We celebrate datasets that come pre-packaged, their columns neatly labeled, their contents ready for algorithmic consumption. This is the world of "structured" data, and it’s the gold standard for a reason: it’s easily queried, analyzed, and integrated. But what about the rest? What about the sprawling, chaotic, and deeply human mess of text found in city council meeting minutes, old court depositions, or decades of internal agency newsletters? This is the domain of "unstructured" data, and we’ve long treated it as a problem to be solved, a chaotic wilderness to be tamed with natural language processing algorithms before it can be considered truly useful.

But what if we’ve been looking at it wrong? What if, in our rush to impose order, we are sanitizing the very essence of the record? The true value of a city council transcript isn't just the final vote tally, which can be neatly structured. It's the meandering debate, the heated exchange, the sudden joke that breaks the tension, the local idiom used by a lifelong resident. These are the elements that reveal the culture, the priorities, and the unspoken social contracts of a place and time. A sentiment analysis algorithm might label a passage as "negative," but it will miss the resigned sarcasm of a council member who says, "Well, that’s just peachy," in response to a budget shortfall. The structure we impose can often erase the scent of reality.

Reading the Scraps

The impulse to structure everything is a modern one, born from a desire for computational efficiency. But history, and the human experience it records, is inherently unstructured. Archivists and historians have always known this. They don't just read the official proclamation; they study the marginalia in the drafts, the personal letters between the principals, the diary entries written the night before. These fragments, these unstructured scraps, are where context lives. They are the difference between knowing that a law was passed and understanding why it was passed, who fought for it, and what compromises were made in the dead of night.

When we apply this lens to digital preservation, the value of keeping the "messy" versions becomes clear. An OCR-scanned PDF of a 1980s community newspaper, complete with scanning errors, quirky fonts, and faded photographs, is a richer artifact than a perfectly transcribed plain-text file. The flaws and the format are part of the record. They tell us about the technology of the time and the materiality of the original object. The structure is a useful abstraction, but the unstructured source is the primary source.

This isn't an argument against structure. Clean, queryable data is powerful and necessary for many tasks. It is, however, an argument for a more humble and inclusive approach to preservation. Our goal shouldn't be to eliminate the unstructured, but to preserve it alongside the structured interpretations we create. We need to build archives that honor the raw, unvarnished testimony of our digital and analog pasts, not just the sanitized summaries. The future historian will thank us not for the flawless database of vote counts, but for the digitized recording where they can hear the tremor in a citizen's voice as they plead for their neighborhood—a piece of data no algorithm can fully structure, and none should ever erase.

Notes & further reading

A few pages I came back to while writing this: