The Bias of the Backup: Why Perfect Preservation Stifles Context

In the world of digital preservation, the dominant creed is one of perfect fidelity: capture every bit, maintain every link, ensure the artifact is rendered exactly as it was. We are told to back up obsessively, to use tools that snapshot a website in its pristine, functional state. The goal is a flawless digital taxidermy, where nothing of the original’s form is lost. But what if this pursuit of perfect preservation is, paradoxically, a way of losing something vital? What if the sterile, complete backup erases the very evidence of decay and change that gives a digital object its true historical texture?

Consider the common advice to regularly back up your personal blog. The result is a series of identical, frozen copies, each one a Platonic ideal of the site at a given moment. This seems responsible. Yet, compare this to the experience of stumbling upon an ancient, neglected GeoCities page through the Wayback Machine. You don’t just see the page; you witness its collapse. The images are half-broken, rendered as placeholder icons. The layout is fractured, tables misaligned by missing style sheets. The links are a mix of the active and the eternally pending. This broken state isn’t a failure of the archive; it is the archive. It tells a story that a clean backup cannot: a story of technological senescence, of abandoned standards, of the web’s organic erosion.

The Aesthetics of Loss

Our obsession with perfect bit-for-bit preservation implicitly prioritizes the creator’s intended experience over the historian’s—or the casual observer’s—lived experience of time. It assumes that a missing JPEG is merely an error to be corrected, not a meaningful event in the life of the document. This approach creates a sanitized history, one where the medium’s fragility and the passage of time are airbrushed out. We preserve the song, but delete the crackle of the vinyl.

This has profound implications for open data and public records. When a government agency ‘migrates’ a dataset to a new, clean, standardized format, they often do so with the best intentions: accessibility, usability, longevity. But in scrubbing the old file formats, the peculiar database structures, the idiosyncratic field names, they erase the administrative context. The struggle to parse a poorly-documented, legacy CSV file is not just a technical headache; it is a direct encounter with the constraints, priorities, and bureaucratic culture of the institution that produced it. The ‘messy, unscrubbed dataset’ is valuable not in spite of its flaws, but because of them.

Perhaps, then, we need to challenge the advice that preservation equals perfection. Maybe we should be designing archival systems that consciously capture and present the decay. Not as a bug, but as a feature. This means keeping the 404s alongside the redirects, preserving the broken layouts as readily as the functional ones, and valuing a dataset’s obsolete structure as a key to its origin. It means embracing an archive that shows its age, that lets the digital artifact weather and patina, not one that forever suspends it in amber. For context is not just the content itself, but the visible mark of all the forces that acted upon it. To preserve only the perfect specimen is to collect butterflies with pins, and mistake them for the living, fluttering thing.

Notes & further reading

A few pages I came back to while writing this: