The Man Who Indexed the Sky: How a 19th-Century Astronomer Prefigured Our Data Struggles

Long before the concept of ‘open data’ was coined, a quiet, meticulous revolution was underway at the Harvard College Observatory. There, in the late 1800s, a team of women known as the ‘Harvard Computers’ was undertaking a project of staggering ambition: to catalog and classify every single star in the night sky. Their leader was a figure whose work feels startlingly modern, a patron saint of data integrity and accessible public records—Edward Charles Pickering.

Pickering’s mission was not to make new discoveries himself, but to build the infrastructure for discovery. He understood that the raw observations of the universe—thousands of glass plate photographs—were useless if they couldn’t be read, compared, and referenced. His great work, the Henry Draper Catalogue, was not a telescope but a index. It was an attempt to create a permanent, readable, and universally accessible record of the heavens, a public good for astronomers the world over.

In this, he prefigured our core digital preservation challenges. His ‘data’ was analog, captured on fragile glass plates coated in silver gelatin. His ‘software’ was the human brain, trained to recognize spectral patterns. His ‘format’ was a handwritten annotation in the margin of each plate. The threat of obsolescence was not a dying file format, but a fading emulsion or the loss of the institutional knowledge held by his team of computers.

Pickering’s solution was radical openness and meticulous standardization. He didn’t hoard the plates; he published the catalogs. He didn’t let the data languish in a proprietary system; he created a standardized classification scheme—the OBAFGKM stellar classification still used today—so anyone, anywhere, could speak the same language. He built a system meant to outlive him, ensuring the data remained a living, breathing resource, not a museum piece.

Today, we grapple with preserving terabytes of digital sky surveys, ensuring future scientists can read our FITS files and understand our metadata schemas. We are the modern equivalents of Pickering’s crew, striving to keep the data alive and meaningful. His work stands as a powerful historical testament: preservation is not merely about saving the artifact—the glass plate, the hard drive—but about saving the means to understand it. It is about building the index, defining the standard, and, most importantly, sharing it with everyone who might look up and wonder.

Notes & further reading

A few pages I came back to while writing this: