The Fiction of Finality: Why No Data Snapshot Is Ever Complete
There’s a comforting promise at the heart of many open data and web archiving projects: the promise of the capture. We speak of ‘taking a snapshot’ of a website, ‘preserving a corpus’ of public records, or ‘creating a static copy’ of a dynamic database. The metaphor is powerful and deliberate. It suggests a moment frozen in time, a perfect and total record, as if someone had flipped a switch and everything was saved, just as it was. It’s a fiction, but a necessary one. Without it, the monumental task of digital preservation would feel impossible from the start. Yet, our faith in this finality is the very thing that can undermine the archives we create.
The reality of a data capture is far messier than the snapshot implies. A web crawler, for instance, doesn’t descend upon a site with the omniscience of a camera shutter. It arrives with a set of instructions, a limited amount of time, and a finite capacity. It may ignore dynamic content that loads with user interaction. It might be blocked by a robots.txt file, or fail to follow complex JavaScript. It captures what it is programmed to see, which is never the entirety of the user experience. What we archive is not the website, but a particular, technologically-determined reading of it. The snapshot is a portrait painted by a machine with a specific set of brushes.
This partiality becomes even more pronounced with large-scale public datasets. When a government agency releases a ‘complete’ dataset of, say, business licenses or property transactions, we accept it as a definitive record. But what about the records being amended hours after the snapshot was taken? What about the data that was excluded for privacy or ‘quality’ reasons before the export was even run? The snapshot presents a veneer of objectivity, obscuring the countless human and automated decisions that shaped its contents. It freezes a process, not a perfect object.
This isn’t to say the work is worthless. On the contrary, acknowledging the inherent incompleteness of our archives is what makes them more valuable, not less. It forces us to be better custodians. It pushes us to document the ‘how’ of collection as meticulously as the ‘what’. Instead of presenting a dataset as a final truth, we can present it as evidence of a system’s state at a specific moment, with clear metadata about the tools used to collect it and the known limitations of the process. The archive becomes a conversation starter about methodology, not a period at the end of a sentence.
Ultimately, the pursuit of a complete snapshot is a Sisyphean task. The digital world is a river, not a lake. By fetishizing finality, we risk despairing when we inevitably fail to achieve it, or worse, misleading future historians who take our ‘snapshots’ at face value. The real work of preservation is not about achieving a state of perfection, but about maintaining a faithful, well-documented process. It’s about building a library where every book has a detailed provenance slip, explaining not just its contents, but the journey it took to get to the shelf. In letting go of the fiction of the complete capture, we embrace a more honest and ultimately more resilient form of memory.
Notes & further reading
A few pages I came back to while writing this:
- Palmdale, CA
- The Warden of the Word List: Salvaging Structure from Unruly Text Files
- Pasadena, CA
- The Digital Museum's Deceit: Why We Should Preserve the Noise, Not Just the Signal
- Pomona, CA
- The Forgotten Scribe of the 1890 US Census
- Riverside, CA
- Roseville, CA
- Sacramento, CA
- Salinas, CA
- San Bernardino, CA
- San Diego, CA
- San Francisco, CA