Making the Evidence Visible

The Society’s archive contained decades of volunteer and bulk-imported material. Some references were hidden from editors when a citation contained only a source name or only a URL. Restoring those fields exposed citations in more than 2,700 honoree records and 25 company records so they could be reviewed and corrected.

That was an important first step: an archive cannot audit information that its own editors cannot see. The citation fields were made consistent across collections before the larger review began.

Checking Sources, Not Merely Links

A working URL is not necessarily evidence. The process evaluated each record together with its cited sources and asked whether the source actually supported the claim. It also rechecked suspected dead links carefully enough to distinguish a missing page from a temporary rate limit.

That distinction prevented temporary website refusals from being mistaken for lost scholarship. In one recheck, 1,327 of 3,267 honoree links previously marked as broken were actually available. The failed first pass became a useful control lesson: the checking system itself also had to be tested.

Across the hardware, software, and company collections, 13,476 records received a verified-current source or a clearly labeled Internet Archive copy.

More Than 13,800 Records Improved with an Audit Trail

AI agents helped examine more than 30,000 records and revise more than 13,800. Every edit included attribution and a plain-English reason and could be reversed by a curator. That audit trail was part of the design from the beginning: scale was useful only if a human could inspect what changed and why.

The work went beyond repairing dead URLs. When a cited page was available but did not support the statement attached to it, the agents looked for better evidence. When the evidence remained uncertain, the item was escalated rather than rewritten with unjustified confidence.

Turning a Flat Catalog into a Useful Hierarchy

The hardware catalog had accumulated 42 flat categories and a large “Other” bucket. The revised taxonomy uses 13 top-level families, 104 subtypes, and 37 granular types. More than 11,000 of roughly 16,800 items were recategorized, including 2,791 previously assigned to “Other.” Related classifications were also refined where a single label could not accurately describe a person or work.

The hierarchy made the collection easier to browse, but it also made its weak spots measurable. A large catch-all category is not merely untidy; it is evidence that the catalog no longer describes the material it contains.

Keeping Humans in Charge

The agents operated as research assistants, not autonomous authors. Versioned instructions made the process reproducible; routine cases used lower-cost models; uncertain cases escalated to a stronger model and then to a human curator. Existing human classifications received extra weight, and bulk-imported records retained their unknown-author designation.

Those controls addressed two different risks. Versioned instructions and attributed edits made the process repeatable; human escalation prevented the system from manufacturing certainty where the historical record was ambiguous. The Society could gain the speed of automated research without pretending that automation had become the scholar of record.

A draft AI policy was developed alongside the work to address provenance, quality control, audit trails, and human authority. It is a proposal for consideration — not a policy approved by the Board.

A More Trustworthy and Navigable Archive

The result is stronger evidence, clearer cataloging, and an operating method that makes AI-assisted research reviewable rather than opaque. Readers get a more navigable archive; curators get a record of what changed, why it changed, and how to undo it. The work is publicly documented in the IT History Society changelog; the draft AI policy describes the controls under consideration.