The Unreliable Witness of Perfect Metadata

In our quest to preserve the digital, we have become obsessive cataloguers. The prevailing wisdom is clear: without rich, structured, machine-readable metadata, a digital artifact is a orphan, a ghost in the machine destined to be un-found, un-contextualized, and ultimately, lost. We pour immense resources into creating meticulous digital footprints for every file, every tweet, every web page, believing that the path to true preservation is paved with standardized tags. But what if this very obsession is, in some subtle ways, corrupting the historical record? What if our reliance on perfect metadata is making the digital past less true?

The problem isn’t metadata itself—it’s the illusion of objectivity it creates. A historian sifting through a box of physical letters finds context in the smudged postmark, the idiosyncratic handwriting, the faded ink, the random shopping list scribbled on the back. These are forms of metadata, yes, but they are intrinsic and often accidental. They are traces of the object’s life. Our digital metadata, by contrast, is almost entirely extrinsic and deliberate. It is a story we, the archivists, tell about the object at a specific moment in time. We decide what is important enough to tag, and in doing so, we silently decide what is not. We impose a contemporary taxonomy on the past, creating a filter through which all future discovery must occur.

This system creates a feedback loop of significance. Items with ‘good’ metadata are easily discovered, cited, and preserved, thus justifying the initial effort and reinforcing the hierarchy. The things that don’t fit our current schema—the weird, the ambiguous, the culturally specific, the plain inconvenient—are pushed to the margins. They become the ‘dark data’ of the archive, not because they lack value, but because they lack the correct keywords. We are, in effect, building an archive that is wonderfully searchable but perilously pre-interpreted. The messiness of history, its contingent and chaotic nature, is smoothed over by the orderly fields of a Dublin Core form.

The Seduction of the Query

The greatest danger lies in the seductive power of the search bar. It encourages a form of historical inquiry that skims the surface, that finds exactly what it’s looking for and little else. The serendipitous discovery—the letter tucked inside a book, the unexpected connection made while turning physical pages—is engineered out of the system. We trade the sprawling, uncertain landscape of the un-catalogued attic for the sterile efficiency of the database query. The archive becomes a place for confirming hypotheses, not for stumbling upon the truths we didn’t know we needed to find.

Perhaps the most counterintuitive proposal, then, is to deliberately leave some things ‘uncatalogued.’ Or, more realistically, to supplement our rigid taxonomies with methods that capture the accidental and the ambiguous. Could we build archiving tools that log their own quirks and failures? Could we preserve not just the ‘clean’ final version of a website, but the broken links, the missing images, the 404 errors as meaningful artifacts in their own right? The gaps and glitches are data, too. They testify to the fragility of the system and the reality of loss. A perfect metadata record implies a perfectly preserved object, which is a fantasy. An imperfect object, with all its scars and silences, is often a more reliable witness to its own history.

Our goal should not be to construct a flawless index for a graveyard of digital objects. It should be to build a rich, chaotic, and even contradictory ecosystem where the past can retain some of its original voice, not just the one we have chosen to give it. The true archive may not be the one that is easiest to search, but the one that is still capable of surprise.

Notes & further reading

A few pages I came back to while writing this: