The Peril of Perfect Provenance: When Tracking Data’s Origin Obscures Its Present

In the world of digital preservation and open data, provenance is gospel. We are taught, correctly, to cherish metadata that tells us where a dataset came from, who touched it, and when. This chain of custody is the bedrock of trust. It turns a raw number into evidence. But I want to argue a counterintuitive, even heretical, point: our fetish for perfect provenance can sometimes become a tool of exclusion, locking data in a past context and rendering it less useful, less alive, and less equitable in the present.

Consider a municipal dataset—say, property boundaries. The perfectly curated version lives in an archive with impeccable provenance: it was created by the city’s GIS department in 2012, updated after the 2015 annexation, and migrated to a new system in 2020. Every change is logged. Yet, over those years, a community land trust has been maintaining its own, parallel map. They’ve corrected errors the city missed, added informal paths and shared garden plots, and annotated properties with oral histories. Their version is messy. Its provenance is a tangle of community meetings, handwritten notes, and smartphone GPS pins. By the strict standards of archival science, it’s ‘unreliable.’ But by the lived reality of the neighborhood, it is vastly more true.

The dogma of pristine provenance often demands that data be ‘official’ to be credible. This creates a hierarchy where the government dataset, with its clean lineage, is deemed authoritative, while the community dataset is seen as derivative or suspect. We preserve the former with care and consign the latter to oblivion, effectively erasing the corrections and context added by those who use the information most intimately. In our quest to track origin, we freeze a single, sanctioned origin story in place.

This extends to web archiving as well. We capture a page, along with its headers and timestamps (provenance), and seal it in a WARC file. That snapshot is ‘true’ for that millisecond. But what about the story of that page’s evolution? The broken link that was fixed the next day, the typo in a critical policy document that stayed up for a week, the community forum thread that was the *real* source of an idea later presented formally on an ‘authoritative’ page? The perfect provenance of the individual snapshot blinds us to the messy, collaborative, and often socially significant process that created the public record.

I am not arguing for sloppiness or against accountability. Instead, I’m advocating for a broader, more humane definition of what makes data preservable. We need systems that can embrace polyphony—that can hold the city’s map and the community’s map in dialogue, without forcing one to conform to the other’s pedigree. This might mean archiving ‘forked’ datasets, preserving contradictory versions, and developing metadata schemes that capture collaborative and contested origins, not just bureaucratic ones.

Preservation is not just about maintaining a chain of custody back to a sanctioned source. It is about maintaining a chain of meaning forward to a living public. Sometimes, to keep data truly open and readable, we must be willing to let go of the perfect story of where it came from, and make space for the messy, vital story of what it became, and in whose hands.

Notes & further reading

A few pages I came back to while writing this: