AI-Generated Finding Aids: What the Technology Can Actually Do, and Where It Still Needs You

Breaks down "AI-generated finding aids" into six distinct sub-tasks (HTR/OCR, NER, scope-note drafting, arrangement, authority reconciliation) and assesses which are production-ready versus pilot-stage, arguing transcription quality is the foundation everything else inherits.

Leo Team

July 26, 2026

AI-Generated Finding Aids: What the Technology Can Actually Do, and Where It Still Needs You

This is a task-by-task account of AI-generated finding aids — what the term actually covers, which parts are ready for real archival work, and where the failure modes hide. If you manage a backlog and are weighing where automation helps and where it quietly introduces error, this piece is written to help you scope it honestly.

An AI-generated finding aid is not, in current practice, a single button that turns a box of unprocessed manuscripts into a DACS-conformant collection guide. The term covers at least six separable tasks: text acquisition through HTR or OCR, named-entity recognition, scope-and-content summarization, biographical or administrative history drafting, arrangement suggestions, and entity reconciliation against authorities like VIAF and LCNAF. Some of these are deployed at scale today. Others are pilots. The fully automated, unattended pipeline is aspirational, and no authoritative source documents one producing a standards-conformant finding aid without human review. The near-universal recommendation across the profession is AI-assisted description with a human in the loop — not AI-generated description left to run alone.

What follows takes that pipeline apart task by task, so you can see which parts are ready for real work, which need careful scoping, and where the failure modes hide. The framing matters because most disappointment with "AI finding aids" comes from treating a chain of distinct tools as one automatic step.

Why the pressure to automate is real

The motivation is not novelty. It is the backlog. Greene and Meissner's More Product, Less Process argued twenty years ago that traditional item-level processing is simply too slow — that survey-based processing benchmarks exceed available archivist hours by large margins, and that making more collections visible faster is worth trading against exhaustive description. The ARL Special Collections Task Force documented the same pattern sector-wide: large unprocessed and underprocessed backlogs of rare book, manuscript, and archival material, with no realistic path to clearing them at current staffing.

That is the demand-supply gap AI is being asked to close. It is a genuine problem, and it is the right frame for judging any tool: does this actually shift the balance MPLP named, or does it just move the labor from description to verification? For most of the pipeline, the honest answer today is that it moves the labor — which is still valuable, but only if you know where the labor lands. This is part of the broader question of what AI reliably does and doesn't do in cultural heritage, and finding-aid generation is one of the clearer test cases.

The pipeline, task by task

Text acquisition (HTR / OCR) — deployed, and the foundation everything rests on

Every downstream task depends on the text underneath it. For handwritten and historical-print collections, that text does not exist until a transcription tool produces it, and its quality sets a ceiling on everything after. This is the single most important thing to understand about the whole pipeline, because the dependency is not gentle. It is a cliff.

Huynh et al. quantified it precisely: named-entity recognition F-scores drop by about 30 percentage points at 20% character error rate and 50% word error rate — and the degradation is worse for historical and vernacular text. Feed noisy transcription into your NER step and you do not get slightly worse access points; you get access points that have collapsed. The scope note built on top inherits the same corruption, but in fluent prose that hides it.

So the first decision in any finding-aid automation project is not which description tool to buy. It is how you will get research-grade transcription out of the actual hands in your collection. If you are working from handwritten or early-modern printed material, this is where general-purpose OCR built for clean modern type struggles — the long s read as f, ligatures split, typographic abbreviation dropped, uneven inking and show-through read as character evidence. It is also the stage most people underestimate. The handwritten text recognition guide covers how HTR differs from OCR and how to read CER figures honestly; the boundary between them matters most exactly here, where a few percentage points of error cascade into the rest of the pipeline.

Named-entity recognition — a pilot, and language-dependent

Extracting persons, places, and organizations as access points is where much of a finding aid's discovery value lives. It is also where the multilingual reality bites hardest. Results do not transfer across languages the way vendor marketing implies. Provatorova et al. showed that specialized historical models beat contemporary multilingual baselines by a large margin on Dutch historical text; Abadie et al. found the same on 19th-century French. A shared cross-language benchmark is only now emerging, and Liu et al. report that even under the most lenient criteria, the highest F1-score remains below 70% when existing NER systems meet historical text across corpora.

The practical consequence: an English-trained or modern-trained transformer will not simply work on your French notarial records or your German parish books. And the realistic working corpus of most mixed special collections is exactly that — vernacular European languages in the Latin alphabet, not Latin. Latin is one historical language among many. Tool selection should follow the corpus in front of you, not a default. The guide to which languages and scripts HTR can actually read makes the case that real support is a function of script, language, and hand together, not a language count on a marketing page.

Scope-and-content and biographical drafting — the most tempting, the most dangerous

This is the task people reach for first, because it looks like exactly what large language models do well: write fluent, plausible descriptive prose. NC State University Libraries ran a careful single-institution test using ChatGPT to draft a collection guide and compared it against a human-written baseline. It is a qualitative case study, not a multi-institution benchmark, and it should be read as one data point rather than a green light.

The danger is specific and worth stating plainly. LLM-drafted scope notes and biographical histories are fluent, and fluency is precisely what makes their errors hard to catch. A garbled OCR line announces itself; a confidently fabricated date, relationship, or provenance detail reads exactly like a correct one. No archival-description-specific hallucination rate has been published, but the nearest transferable figure is sobering: Chelli et al. measured reference hallucination at 39.6% for GPT-3.5 and 28.6% for GPT-4 in systematic-review writing — a citation-style task structurally similar to composing a biographical history with dates and attributions. Treat that as an upper-bound warning, not a measured archival rate, but do not treat it as irrelevant.

The discipline here is the same one that governs any machine-assisted reading: keep the base record and the interpretation separate, and verify the interpretation against the source. Fluent-but-wrong output is the defining failure mode of generative systems on historical material, and a scope note is interpretation layered on top of a transcription — two places for plausible error to enter, not one.

Arrangement suggestions and classification — assistive only

Automatic document classification and arrangement suggestions can surface patterns across a large series faster than a person skimming boxes. But arrangement is an intellectual act rooted in provenance and original order, and a suggestion is not a decision. Use these outputs to orient, not to structure. Nothing in the current literature supports handing arrangement to a model unattended.

Entity reconciliation against authorities — partly mechanical, wholly checkable

Matching extracted names against VIAF, LCNAF, or Wikidata is the most tractable step, because the target is a controlled vocabulary and the match can be verified against an authority record. It is also where AI-introduced bias in access points and subject terms is most likely to slip in unexamined, and the governance conversation about how to flag and audit that is still early. The Cambridge Forum discussion of AI provenance is theoretical, and the SAA Generative AI Statement foregrounds transparency without specifying a binding recording convention. There is, as yet, no standard EAD or RiC field that tells a downstream reader "this access point was AI-suggested." Until there is, that transparency is a local decision you have to make deliberately.

Why "LLM prose alone is a finding aid" is the costliest misconception

It is worth being blunt about the standards gap, because it is where automation promises most often break. A finding aid is a structured description. DACS expects defined elements — biographical or administrative history, scope-and-content, container list, access points — with crosswalks to EAD, EAC-CPF, MARC 21, and RDA. EAD3 requires encoding against a defined tag library. Records in Contexts expects typed relations in a graph ontology.

An LLM produces unstructured or semi-structured prose. It is a draft until an archivist maps it to a standard and authorizes it against authorities. The gap between "readable paragraph about a collection" and "conformant, encoded, authorized finding aid" is exactly the gap where professional judgment lives — and it is not closing on its own. Vendor claims deserve the same scrutiny: Preservica's 2026 announcement describes image-side person, place, and object detection and metadata assistance, not DACS/EAD finding-aid generation, and ArchivesSpace, AtoM, and CollectiveAccess ship no documented AI-description module in their public release notes. Vet every claim against published features, not roadmap language.

Where a purpose-built transcription tool fits in this pipeline

The task-by-task view has a clear implication: fix the foundation first. Because NER, summarization, and biographical drafting all inherit whatever error the transcription step leaves behind, the highest-leverage decision in the whole chain is the quality and integrity of the text you start from. This is the stage — reading the actual hands and historical print in the collection, accurately and faithfully — where a specialist tool earns its place, and where Leo is built to work.

Leo's transcription model, ATR-1, reads Latin-script manuscripts and printed matter from roughly the last five centuries — whatever the language on the page, with performance strongest in English and strong across French, German, Spanish, Italian, Dutch, and other languages written in that alphabet. It reads secretary hand, Gothic cursive, and italic, as well as historical print where conventional OCR falters. Two design choices matter for finding-aid work specifically. First, it transcribes what is on the page rather than smoothing it into modern, plausible prose — the long s stays a long s, an archaic spelling survives, a marginal addition is preserved — so the text you feed downstream is the record, not a normalization of it. Second, it errs recoverably: output that shows failure patterns is detected, retried, and refunded rather than passed through, and the errors that do get through are wrong characters or words you can check against the image beside the text, not fluent fabrications. That distinction is the whole safety argument for the foundation layer, because a hallucinated transcription poisons every task built on it.

On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, ATR-1 scored roughly 5% character error rate at release — 61% fewer errors than the next-best model tested (Transkribus/Text Titan I ~13%, Claude Opus ~23.3%, Gemini 2.5 Pro ~24.8%, GPT-4.1 ~56.7%), per the published benchmark. Given the roughly 30-point NER collapse Huynh et al. document at high CER, the value of starting the pipeline lower on that curve is not cosmetic. Leo also carries structured, Dublin-Core-adjacent metadata fields per document, global fuzzy search across transcriptions, and TEI XML export — so the transcribed collection arrives as searchable, described text ready for whatever description work comes next, rather than as loose output you have to wrangle. It does not produce EAD3 or RiC, and it does not claim to; that encoding and authorization remains, correctly, the archivist's. For a fuller picture of how the transcription stage sits alongside capture, metadata, and delivery, the planning guide to making archives searchable maps the surrounding pipeline.

A working stance for archivists

The honest position, supported across the current literature, is neither refusal nor surrender. Scope AI to specific sub-tasks. Budget for transcription quality first, because everything downstream inherits it. Select language tools by the corpus actually in front of you, not by a marketing count. Treat every generated scope note and biographical history as a draft to be verified against the source and encoded by hand. Record, in whatever local convention you can, which elements were AI-assisted, since the standards bodies have not yet given you a binding one.

None of this replaces the descriptive judgment at the center of the profession — the reading of provenance, the weighing of original order, the decision about what a collection is about. AI can clear the mechanical undergrowth around that judgment and make more of the backlog visible faster, which is exactly what MPLP asked for. What it cannot do is make the judgment for you, and the finding aids that hold up will be the ones where a person still made it.

Frequently Asked Questions

What are AI-generated finding aids and can they replace archivists?

AI-generated finding aids refers to using AI across a chain of separable tasks — text acquisition through HTR or OCR, named-entity recognition, scope-and-content summarization, biographical or administrative history drafting, arrangement suggestions, and entity reconciliation against authorities like VIAF and LCNAF. No authoritative source documents a fully automated pipeline producing a standards-conformant finding aid without human review. The profession's near-universal recommendation is AI-assisted description with a human in the loop, not AI-generated description left to run alone. AI can clear mechanical work and make backlogs visible faster, but the descriptive judgment — reading provenance, weighing original order — remains the archivist's.

Why does transcription quality matter so much for AI finding aids?

Transcription quality sets a ceiling on every downstream task, because named-entity recognition, summarization, and biographical drafting all inherit whatever error the text step leaves behind. The dependency is a cliff, not a gentle slope: named-entity recognition F-scores drop by about 30 percentage points at 20% character error rate and 50% word error rate, and the degradation is worse for historical and vernacular text. Feed noisy transcription into your NER step and access points collapse; the scope note built on top inherits the same corruption, hidden in fluent prose. The first decision in any finding-aid automation project is how to get research-grade transcription from the actual hands in your collection.

Can ChatGPT write a scope-and-content note or biographical history for an archival collection?

ChatGPT can draft fluent, plausible descriptive prose, but that fluency is exactly what makes its errors dangerous. A garbled transcription announces itself; a confidently fabricated date, relationship, or provenance detail reads exactly like a correct one. NC State University Libraries tested drafting a collection guide with ChatGPT as a single-institution qualitative case study, not a benchmark. No archival-description-specific hallucination rate has been published, but reference hallucination has been measured at 39.6% for GPT-3.5 and 28.6% for GPT-4 in structurally similar citation tasks. Treat every generated scope note or biographical history as a draft to verify against the source and encode by hand.

Why can't an LLM produce a DACS or EAD3 conformant finding aid on its own?

An LLM produces unstructured or semi-structured prose, while a finding aid is a structured description with defined elements — biographical or administrative history, scope-and-content, container list, access points — encoded against standards. DACS expects those elements with crosswalks to EAD, EAC-CPF, MARC 21, and RDA; EAD3 requires encoding against a defined tag library; Records in Contexts expects typed relations in a graph ontology. A readable paragraph about a collection is a draft until an archivist maps it to a standard and authorizes it against authorities. That gap — between prose and conformant, encoded, authorized description — is where professional judgment lives, and it is not closing on its own.

Does named-entity recognition work the same across different historical languages?

No — named-entity recognition results do not transfer across languages the way vendor marketing implies, and performance is strongly language-dependent. Specialized historical models beat contemporary multilingual baselines by a large margin on Dutch historical text, and the same holds for 19th-century French. Even under the most lenient criteria, the highest F1-score across historical corpora remains below 70%. An English-trained or modern-trained transformer will not simply work on French notarial records or German parish books. The realistic working corpus of most mixed special collections is vernacular European languages in the Latin alphabet, so tool selection should follow the corpus in front of you, not a default.

© 2026 Leo Technologies Limited. All rights reserved