The Data Tombstone: What an Obituary Teaches Us About Digital Preservation

In the quiet corners of old newspapers and local history websites, obituaries do more than announce a passing. They are dense, structured packets of context. A name, a date of birth, a date of death, surviving relatives, perhaps a profession or a hobby. For genealogists and historians, this structure is a godsend. It’s a model of clarity. And it’s precisely this model, this idea of perfect, self-contained metadata, that has become a kind of received wisdom in digital preservation. We are taught to believe that if we can just tag our data correctly, if we can create a perfect digital tombstone for every file, its future is secured.

The allure is obvious. A file named ‘Meeting_Minutes_2024_03_21.pdf’ is infinitely more useful than ‘scan00543_final_v2.pdf’. We fill our preservation systems with fields for creator, creation date, subject, and format. We build these meticulous digital headstones, believing that the information carved upon them will tell future users everything they need to know. We are, in effect, creating archives of perfectly labeled tombstones. But an obituary, for all its utility, is not a life. It’s a summary, a sanitized conclusion. It captures the endpoints and the headlines, but it silences the messy, vital narrative that existed in between.

Our obsession with pristine metadata creates what I call the ‘Data Tombstone’ problem. We preserve the document, and we preserve the descriptive placard, but we often fail to preserve the ecosystem that gave it meaning. The ‘Meeting_Minutes_2024_03_21.pdf’ might be perfectly preserved, but the sprawling email thread that led to the meeting, the shared document drive where the agenda was drafted, the instant messenger banter that clarified a key point—these are almost always lost. The context is ephemeral. We save the official record but lose the lived experience of its creation.

This is more than just a loss of color; it’s a loss of understanding. Future historians might know that a decision was made on a certain date, but without the context of the debate, the alternative proposals, the social pressures, and the unspoken assumptions, their interpretation of that decision will be stunted. They will have the tombstone, but not the biography. The obsession with the tombstone can lead us to privilege certain types of ‘clean’ data—final reports, published articles, official statements—over the ‘dirty’ data of process and collaboration, which is often where the real story lies.

The challenge, then, for those of us invested in open data and digital preservation, is to move beyond the tombstone model. We must acknowledge that while good metadata is crucial, it is also insufficient. The real work lies in developing practices and tools that can capture and preserve context. This might mean archiving not just the final PDF, but the version history of the Google Doc it came from. It might mean encouraging the preservation of communication channels alongside their formal outputs. It’s a vastly more complex undertaking, one that embraces the messiness of how knowledge is actually made. We need to stop trying to write perfect obituaries for our data and start thinking about how to preserve its lively, complicated biography.

Notes & further reading

A few pages I came back to while writing this: