The Argument for Incompleteness: Why Partial Public Records Are Often More Ethical

The guiding light of open data and public records has long been a simple, powerful one: more is better. The drive is towards comprehensiveness, completeness, the total capture of information to ensure accountability and fuel innovation. It’s an ethos that has served us well, pulling back curtains on government operations and corporate influence. But in our relentless pursuit of the ‘complete’ dataset, we risk creating a blunt instrument that, in its totality, can become a tool of quiet but profound harm.

Consider a common, almost sacred piece of advice in this field: that we must fight to archive and publish every scrap of a public record, redacting only what is legally mandated. This impulse is born from a deep-seated and often justified distrust of authorities who might hide their misdeeds behind selective disclosure. Yet, this absolutist approach ignores a more nuanced reality—that within the vast administrative machinery of the state, there exists not only the powerful, but also the vulnerable. They are citizens who interact with social services, individuals whose personal tragedies or medical emergencies become cold lines in a spreadsheet, people whose inclusion in a ‘complete’ public dataset does not serve the public interest, but rather subjects them to a new form of digital exposure they never consented to.

When Totality Becomes a Travesty

We must ask: who does the ‘complete’ record actually serve when it includes, for example, the names and addresses of every recipient of a pandemic relief grant, or the detailed case notes of individuals in public housing disputes? To the researcher or journalist, it’s a richer dataset. To a data mining company, it’s a lucrative source of profiling information. But to the individuals contained within those rows, it is a violation of a reasonable expectation of privacy within a necessary public interaction. Their data becomes a public spectacle not because they sought the spotlight, but because they sought help.

This isn't an argument for secrecy, but for a more sophisticated, humane definition of ‘open.’ It challenges the common advice by proposing that ethical openness sometimes requires deliberate, principled incompleteness. It means designing systems that publish aggregate trends, statistical models, and process audits—the true mechanisms of accountability—while automatically filtering out the granular, personally identifying details of private citizens who are not public figures. The integrity of the record is preserved in its ability to answer the core civic question: ‘Is the system functioning justly?’ It does not require the sacrificial offering of every individual’s private context to the public web.

Preservationists and open data advocates rightfully fear the slippery slope. Who gets to decide what is filtered? The answer must be a transparent, participatory process—a protocol not of redaction, but of thoughtful omission designed into the system from the start. It shifts the burden from justifying why something should be *removed* from a complete dump, to justifying why a specific piece of personal data needs to be *included* in a public archive. This is a harder, more complex path than the brute-force ‘release everything’ model. It accepts that the most perfect technical copy can be the most flawed ethical artifact. In this light, an incomplete record is not a failure of preservation, but a successful act of civic care.

Notes & further reading

A few pages I came back to while writing this: