The Unwritten Glossary: On the Ghosts of Context in Public Records

We often speak of open data as if it were a collection of pristine, self-evident facts, waiting in a digital quarry to be mined. The promise is that by simply making a dataset public, its truth becomes accessible to all. But what happens when the data arrives without its dictionary, when the key to its own language has been lost?

I recently spent an afternoon with a beautifully preserved municipal dataset from the early 2000s, a spreadsheet tracking maintenance requests for public parks. The columns were crisp, the dates were precise, and the work-order numbers marched in perfect sequence. It was a model of digital preservation. And yet, it was also largely incomprehensible. The ‘Status’ column was a cryptogram: ‘CMP,’ ‘DEF,’ ‘AWT-R,’ ‘AWT-A.’ The ‘Priority’ field used numbers, but the scale was a mystery. Was 1 the highest urgency, or was it 5? The data was perfectly preserved, but the context required to read it had evaporated.

The Silent Language of Bureaucracy

Every office, every department, every era develops its own shorthand. These codes and classifications are the living language of daily work, so obvious to the people using them that writing them down feels as unnecessary as defining the word ‘the.’ ‘AWT-R’ likely meant ‘Awaiting Review’ and ‘AWT-A’ probably stood for ‘Awaiting Approval.’ But these are guesses. The civil servant who knew this lexicon by heart has moved on; the internal memo that first established these codes has been deleted; the binder on the shelf that held the master key has been recycled.

This is the ghost in the machine of public records: the unwritten glossary. We preserve the output but lose the schema of understanding. The data is open, but its meaning is locked away. It becomes a kind of digital artifact whose cultural significance we can only infer, like trying to understand a ancient ritual from its tools alone.

This presents a subtle but profound challenge for web archivists and digital preservationists. Our task is not merely to capture the ones and zeroes, but to capture as much of the ecosystem of meaning as possible. It’s the metadata about the metadata. It’s saving the ‘README’ file, the internal wiki page, the training manual, the email thread where someone asks, “What does DEF stand for again?” These are the fragile, often overlooked documents that give primary data its voice.

The next time you encounter a pristine but cryptic public dataset, remember that its full story isn’t just in the cells of the spreadsheet. It’s in the quiet, human agreements that once gave those cells purpose. Preserving data is not just an act of technical salvation, but one of cultural translation. We must strive to archive not only the answers, but the language in which the questions were first asked.

Notes & further reading

A few pages I came back to while writing this: