The Decaying Clock: Can We Archive a Website That Changes Every Second?
I recently found myself staring at a live data dashboard for a global weather monitoring system. Wind speeds, barometric pressure, sea temperatures—every number on the screen was a live feed, updating not by the minute, but by the second. My usual archivist instinct kicked in: this is important public data, it should be preserved. But then I hesitated. What, exactly, would I be preserving? A single, fleeting moment in a relentless, endless stream of data. It felt less like trying to capture a document and more like trying to bottle a river.
This is the unique challenge of archiving real-time data streams and highly dynamic websites. We’ve gotten good at taking static snapshots of web pages. The process for a news article or a government report is relatively straightforward. But what about the website that is, by its very nature, a perpetual state of becoming? The stock ticker, the live traffic map, the continuously updating election results page—these are not fixed entities. They are processes made visible. To capture them in a single snapshot is to fundamentally misunderstand what they are, like taking a single frame from a movie and calling it the entire film.
The Half-Life of a Live Moment
The value of a live data feed is its immediacy. The moment it is captured and frozen in an archive, that value begins to decay. An archived screenshot of a live election map from 8:03 PM on election night is not a record of the election’s outcome; it’s a record of what the algorithms projected at that specific, now-historic, micro-moment. Without the context of the seconds that came before and after, its meaning is fragile, almost archaeological. Future historians would need to piece together thousands of these snapshots to understand the narrative flow, a task far more complex than reading a finalized, official report published the next day.
This presents a philosophical problem for web archivists. Is our goal to preserve the artifact, or the experience? For a static document, they are one and the same. For a live feed, they are irrevocably separate. The artifact (the screenshot) is a pale ghost of the experience (watching the numbers dance in real-time). This forces us to ask a deeper question: when we deem such a stream worthy of preservation, what are we really trying to save? Is it the raw data points themselves, which might be better preserved in a structured database dump? Or is it the public interface, the way the data was presented and consumed by millions of people at a critical moment?
Some institutions are experimenting with "temporal sampling"—capturing high-frequency dynamic content at regular, short intervals. This creates a flicker-book effect of the evolving data. Yet, this approach generates a colossal amount of data and still only approximates the lived experience. It captures the 'what' but loses the essential 'now-ness' that gave the original its power and context.
Perhaps the answer isn't to find a perfect preservation method, but to redefine our expectations. Maybe we must accept that some digital phenomena are inherently ephemeral, and that our archives will only ever hold their echoes. The real-time web is a performance, not a sculpture. We can preserve the script and some photographs of the stage, but the live performance itself disappears into memory the moment it ends. And in acknowledging that intrinsic loss, we might better appreciate the fragile, flowing nature of the present moment we are so desperately trying to pin down.
Notes & further reading
A few pages I came back to while writing this:
- Aurora, IL
- The Unwritten Ledger: Why Some Public Records Are Never Created
- Chicago, IL
- The Conflicted Archivist: To Hoard the Web or To Curate It
- Joliet, IL
- The Unassuming Archive of the Shopping List
- Rockford, IL
- Indianapolis, IN
- Kansas City, KS
- Olathe, KS
- Overland Park, KS
- Topeka, KS
- Lexington, KY