The Unbearable Weight of Perfect Metadata
In the world of digital preservation, we have a sacred mantra: metadata is everything. We are taught that without pristine, granular, and meticulously structured metadata, our archived data is little more than a digital landfill—a chaotic jumble of bits without context or meaning. This pursuit of perfect description has become an article of faith, an unassailable good. But what if this obsession is not just a practical challenge, but a philosophical error? What if, in our quest to perfectly describe the container, we are ensuring the contents inside are never seen?
The common advice is to describe everything. Tag, categorize, and cross-reference until every digital object is enmeshed in a web of contextual data so strong it can never be lost. This is the archivist’s ideal. Yet, this process is incredibly labor-intensive, often requiring specialized knowledge and a level of detail that borders on the obsessive. The result is that we create magnificent, exquisitely described archives of a tiny fraction of what actually exists. We preserve the ‘important’ things perfectly and lose the vast, messy, human-scale everything else. We are building digital museums when we should be aiming for digital wilderness preserves.
This creates a perverse incentive structure. The immense effort required for ‘proper’ metadata becomes a gatekeeper, deciding what is worthy of preservation at all. A city council’s PDF meeting minutes might get the full treatment, but the sprawling, link-rotten community forum where residents actually debated the issues is left to die because its chaotic structure defies easy cataloging. We privilege what we can easily describe over what might actually be valuable. We are, in effect, preserving for the convenience of future librarians, not for the curiosity of future humans.
Perhaps a more radical, and ultimately more humane, approach is to intentionally embrace ‘good enough’ metadata. Instead of trying to describe every tree, we simply mark the perimeter of the forest. We capture the data first, at scale, and trust that future tools and future users will have their own methods of making sense of it. The value is in the sheer existence of the data, not in our contemporary, and inevitably flawed, interpretation of it. The goal shifts from creating a perfectly curated collection to creating a massive, searchable, and yes, sometimes messy, corpus of the past.
This isn’t an argument for carelessness. It is an argument for triage. It suggests that the preservation community’s most precious resource is not storage space, but time. By accepting metadata that is functional rather than flawless, we can use that time to save more of the digital record itself. We must stop letting the perfect description be the enemy of the good enough capture. Sometimes, the most important metadata is simply the knowledge that something was there, waiting for a future mind to ask the right question and find it.
Notes & further reading
A few pages I came back to while writing this:
- Sterling Heights, MI
- The Man Who Indexed the Sky: How a 19th-Century Astronomer Prefigured Our Data Struggles
- Warren, MI
- The Box in the Attic: On Finding My Grandfather's Forgotten Archive
- Saint Paul, MN
- The Unwritten Epilogue: On the Silence of Concluded Datasets
- Springfield, MO
- St Louis, MO
- Jackson, MS
- Cary, NC
- Charlotte, NC
- Fayetteville, NC
- Greensboro, NC