The Humble Semicolon: A Typographic Fossil in the Data Stream
For years, my data work involved meticulously parsing CSV files. The comma-separated values format is the universal, if slightly grimy, currency of open data. It’s plain, it’s legible, and it promises a frictionless flow from database to analysis. Yet, I always felt a tiny, instinctual dread when downloading a new dataset, and it hinged on a single character: the semicolon. Would this file, published by a municipal office or a research body, use commas or semicolons as its delimiter? That one typographic choice, seemingly trivial, became a small but potent artifact of digital preservation and geopolitical data habits.
The semicolon-delimited CSV is not a different format; it’s a dialect. Its prevalence is a quiet map of influence, tracing back to European locales where the comma is already employed as the decimal separator. In much of continental Europe, you write 3,14 for pi. A standard CSV would render that number as two separate columns: “3” and “14”. The semicolon emerges as a pragmatic fix, a cultural workaround baked into the very syntax of data exchange. When you encounter a file from a German environmental agency or a French statistical institute, you aren’t just getting numbers. You’re getting a fragment of their notation system, a digital echo of how they write.
The Ghost in the Spreadsheet
This is where preservation gets interesting. A CSV file is, in theory, perfectly preserved data. The bits are intact, the text is legible. But without the context of its delimiter—a piece of metadata often absent from the file itself—the data becomes garbled. The preservation isn’t of the file alone, but of the unspoken rule that governed its creation. This rule is rarely written in a README; it’s assumed knowledge, a communal practice. Over time, as datasets are migrated, republished, or handed off to international collaborators, that assumption can fade. A semicolon-delimited file from a 2008 EU project, uploaded to a new open data portal a decade later, might silently break a researcher’s script on another continent. The data is saved, but its usability is decaying.
We tend to think of digital fossils as grand things: a GeoCities page, a Flash animation, a proprietary database format. But the humblest fossils are the most revealing. The stubborn persistence of the semicolon in data streams is a fossil of regional convention, a tiny monument to the fact that our global data infrastructure is built atop a thousand local realities. It’s a testament to the fact that true openness isn’t just about publishing a file; it’s about preserving the grammar needed to read it. Every time my parser fails on a semicolon, I’m not just facing a technical hiccup. I’m being politely, if frustratingly, reminded that data is never just raw; it’s cultured, shaped by the hands that prepared it and the places they call home. In that tiny glyph, I see the whole messy, beautiful challenge of building a web of readable public records—a project forever negotiating between one universal standard and a world of particular, persistent commas.
Notes & further reading
A few pages I came back to while writing this:
- Port St Lucie, FL
- The January Reset and the Ghosts in Your Permissions
- Tallahassee, FL
- The Unstoppable Link Rot of Supreme Court Citations: A Critique of the 'Official PDF' Guarantee
- Tampa, FL
- The Art of the Wayback Query: Unearthing a Single Page's Many Pasts
- Augusta, GA
- Columbus, GA
- Savannah, GA
- Honolulu, HI
- Cedar Rapids, IA
- Des Moines, IA
- Boise, ID