The Inevitable Glitch: When Corrupted Data Holds a Deeper Truth
We tend to think of digital preservation as a quest for perfect fidelity. The goal is to capture a pristine, unaltered copy of a file, a website, or a dataset, and then lock it away in a virtual vault, safe from the ravages of bit rot and obsolescence. A corrupted file, in this framework, is a failure. It’s a ‘bad’ copy, destined for the digital recycle bin, an anomaly to be scrubbed from the archival record. But what if we’re missing something? What if the corruption itself is a kind of record?
Consider the JPEG. It’s a format riddled with compromises, a ‘lossy’ compression designed to sacrifice perfect accuracy for the sake of small file sizes. Sometimes, these files become corrupted. A sector on a hard drive fails, a transfer over a shaky network connection drops a few packets, and suddenly, a photograph is sliced with garish colors, or a face is blocky and distorted. Our immediate reaction is to seek out the original, the clean version. But the corrupted version is a unique artifact. It is a document of a specific moment of failure in a specific technological system. It tells a story that the perfect image cannot.
This glitch is a fingerprint of the medium itself. It reveals the underlying structure of the data—how the compression algorithm tiles the image, how the file is structured on the disk. The corruption is not random noise; it is a pattern dictated by the very rules of the digital system that created it. In a way, a severely corrupted file tells you more about the file format and the system that mangled it than a pristine file ever could. The pristine file is opaque; it works as intended. The corrupted file is forced to reveal its inner workings.
This perspective extends beyond images to public records and open data. Imagine a massive CSV file of municipal spending, a cornerstone of government transparency. In its perfect state, it is a clean grid of numbers and categories. But if a scripting error during an export or a storage error on a public server introduces a few malformed rows, we see it as a problem to be corrected. However, that flawed file is evidence. It’s a snapshot of the data pipeline at a moment of breakdown. It can reveal the fragility of the systems that manage our public information, the points where human or machine error can intrude. The error isn’t just a mistake to be erased; it’s a diagnostic tool.
This isn’t to say we should abandon the pursuit of clean data. For practical use, we need accuracy. But for the historical and diagnostic mission of preservation, perhaps we should be more hesitant to discard our failures. A digital archive that only contains ‘perfect’ copies presents a sanitized, frictionless view of technological history. It ignores the reality that systems fail, data degrades, and errors are an inherent part of the digital landscape. By preserving the glitch alongside the intended artifact, we preserve a more honest and complete story. We acknowledge that the path of preservation is not a straight line, but one marked by detours, dead ends, and unexpected, illuminating accidents.
So, the next time you encounter a file that’s been scrambled by time or error, before you hit delete, take a moment. Look at the patterns in the chaos. You might be looking at a more truthful representation of its journey than the file itself ever intended to show.
Notes & further reading
A few pages I came back to while writing this:
- a helpful reference
- The Ghost in the Ledger: When Public Records Are Deleted, But Not Forgotten
- a place-by-place guide
- The Meticulous Stream: Archive-It and Conifer's Divergent Paths Through Time
- a local resource
- The Unassuming USB Drive: A Relic in the Age of the Cloud
- a regional guide
- a useful directory
- one area's overview
- a practical rundown
- a nearby resource
- New York
- Montana