The Unwritten Contract: Who Archivizes the Archivists?

We trust web archives to be our collective memory. When a government website changes policy, when a corporation edits its terms of service overnight, when a blog post vanishes into the ether, we point to the Wayback Machine as the ultimate arbiter of what was. It’s a testament to the success of these institutions that we often speak of their captures as immutable snapshots, digital facts frozen in time. But this faith rests on an unexamined premise: that the process of archiving itself is neutral, a simple act of transcription from the live web to the archive.

This is the received wisdom I want to challenge. The act of preservation is not a passive one. It is an act of creation, and with every creation comes a creator’s bias. We celebrate the archivist as a hero saving data from the abyss, but we rarely ask what lens they look through. The selection of what to archive, the frequency of the crawl, the technical decisions about how to render a complex, interactive page into a flat, static file—these are not neutral technicalities. They are curatorial acts of immense power. An archive that only captures a homepage once a year tells a different story than one that captures a news site every hour during a crisis. An archive that cannot properly render JavaScript-dependent content effectively erases entire functionalities, creating a ghost of a website that never truly existed in that form.

This becomes a profound issue of accountability. If a web archive is presented as evidence in a court case, or used by a historian to understand a pivotal moment, the integrity of that evidence is only as strong as the undocumented decisions made by the archiving software and its human operators. What was excluded? What was broken during the capture? The "original" website is gone, and what we have is the archive’s interpretation of it. The archivist, in this sense, becomes a ghostwriter of history, their choices and limitations indelibly etched into the record we treat as objective.

The Silence in the Metadata

Part of the problem is epistemic. We lack a robust framework for expressing the fragility of a capture. A 404 error is clear; a subtly broken interactive map, or a stylesheet that failed to load, is often silent. The metadata accompanying an archived page might tell us the timestamp of the capture, but it rarely tells us the confidence level of the render, or what alternative capture methods were attempted and failed. We are presented with the final product, its seams carefully hidden, fostering an illusion of comprehensiveness.

This is not a call to abandon web archiving. It is a call to renegotiate our unwritten contract with the archivists. We must move from seeing archives as perfect libraries to understanding them as fallible, human-made collections. This means demanding greater transparency in their methodologies. It means developing richer, more expressive ways to document what a capture attempt missed or mangled. It means acknowledging that preserving the digital world is not like pressing "pause" on a VCR; it is an ongoing, interpretive process fraught with the same subjective challenges as any other form of historical preservation.

For the archive to truly serve as a guardian of truth, it must also archive a record of itself—its own choices, its own failures, its own blind spots. The final, most crucial layer of preservation may not be the data we set out to save, but the context of how we saved it. Because if we don’t archivize the archivists, we risk building a palace of memory on foundations we never thought to inspect.

Notes & further reading

A few pages I came back to while writing this: