Building a Historical Research Workflow: From Archive Photo to Citable Source

Overview of the five-stage individual-researcher workflow (capture, organise, transcribe, analyse, cite) turning archival manuscript photos into citable sources, emphasizing how each stage constrains the next and where source fidelity must be protected.

Leo Team

July 22, 2026

Building a Historical Research Workflow: From Archive Photo to Citable Source

This is a practical account of the historical research workflow — the five stages that turn a manuscript photographed in an archive into a source you can cite with confidence. It matters because each stage constrains the next: a poor photograph limits transcription, an untraceable transcription cannot be cited, and analysis is only ever as good as the text beneath it. Here is how the stages fit together, and where fidelity to the source has to be defended.

A historical research workflow is the end-to-end process that turns a manuscript in an archive into a source you can cite with confidence. It has five stages — capture, organise, transcribe, analyse, and cite — and each one constrains the next. A badly lit photograph limits transcription. A transcription without provenance cannot be cited. Analysis is only ever as good as the text beneath it. The stage that stalls most individual researchers is transcription, where a difficult hand can hold up months of work. What follows walks through all five, names the tools that fit each, and marks the points where fidelity to the source has to be defended.

The workflow is the same whether you are reading Elizabethan wills, French notarial minutes, or German parish registers. What changes is the hand and the language; the discipline does not. This is the individual researcher's version of the broader historical research workflow — not the institutional digitisation pipeline, but the process a single scholar runs between a reading-room visit and a footnote.

Stage one: capture

Everything downstream inherits the quality of the image. If the photograph is soft, skewed, or unevenly lit, no transcription method — human or machine — will fully recover what the shadow swallowed.

Institutions capture to measurable standards. The US federal benchmark, FADGI's Technical Guidelines for Digitizing Cultural Heritage Materials, defines its 4-star tier as native 600 ppi, 16-bit greyscale or 48-bit colour, an embedded ICC profile, and ISO 19264 conformance. The Dutch Metamorfoze Preservation Imaging Guidelines set a colour-fidelity target of ΔE2000 and a noise tolerance of L≤1.6 at the Full tier.

Most individual researchers will never hit those numbers, and for a working transcription they do not need to. What they need is legibility and faithful geometry: the page in focus, filling the frame, shot square-on under even light, saved as a high-resolution file rather than a phone screenshot. The practical craft of doing this in a reading room — handheld capture, lighting, focus, and handling fragile originals without damage — is worth getting right the first time, because reshoots usually mean another trip. A field method for photographing archival documents covers that ground in detail.

One point deserves emphasis. No published study cleanly isolates how much transcription error comes from capture quality versus document condition versus the reading method itself. The vendor standards argue for high capture quality without quantifying the payoff. Treat good capture as cheap insurance rather than a guarantee.

Stage two: organise

A few hundred images become unmanageable fast. The organise stage gives each image a stable identity — a shelfmark, a folio number, a collection — so that a reading can later be traced back to its exact source.

This is what dedicated tools do well. Tropy, free and open-source from the Roy Rosenzweig Center, manages research photographs, supports metadata templates, and exports structured description. Omeka S publishes exhibits with image-delivery support; DEVONthink indexes and OCRs PDFs. Each of these is genuinely useful at what it does.

But there is a common misconception worth naming: organising your photos is not the same as building a corpus. Tropy and its peers produce metadata-rich descriptions of containers and items. They do not read the images. A well-tagged folder of 400 manuscript photographs is still 400 pictures you cannot search the text of — you can find the image, but not the word inside it. Transcription is a separate stage, and organisation does not stand in for it. If you are weighing where an organiser stops and a transcription engine begins, the comparison of Tropy and Leo draws that line precisely.

The minimum metadata to record at this stage is the metadata you will need to cite: repository, collection, shelfmark or call number, folio or page, and a reference to the specific image. Capture it now, while the source is in front of you, rather than reconstructing it later from memory.

Stage three: transcribe

This is the contested stage — the one where the workflow most often stalls, and the one where the choice of method matters most. Transcription converts images into searchable, citable text. Everything after it depends on getting the text right, because errors here propagate into analysis and, eventually, into published work.

There are five broad approaches, and they are not interchangeable.

The methods, briefly

Manual transcription is slow — roughly ten to sixty minutes per page depending on the hand — but it produces authoritative ground truth. It remains the standard against which everything else is measured.

Crowdsourced transcription through platforms like Zooniverse and FromThePage pools consensus from two to five non-experts per image. It is the sensible default for large, homogeneous institutional backlogs, but it depends on volunteer availability and expert-defined conventions, and it does not scale down well to a single researcher's project.

General-purpose OCR — Tesseract, ABBYY FineReader, Google Cloud Vision, Amazon Textract — is engineered for clean modern print and is very good at it. It is the wrong instrument for manuscript hands. ABBYY's engine explicitly limits its handwriting support to handprinted, non-cursive text; Tesseract's documentation says the same. On secretary hand or Kurrent, glyph-classification approaches fail because connected letterforms merge into shapes the template was never trained on. This is the difference between OCR and HTR, and it is worth understanding before you choose a tool — the practical guide to handwritten text recognition for historians sets out the mechanism.

Specialist HTR — Transkribus, eScriptorium, OCR4all — treats a line of text as a sequence rather than a bag of glyphs, which is why it generalises across a writer's variation. It is the reliable path for old hands. The historical catch is that usable accuracy on an idiosyncratic hand has typically required training or fine-tuning a model on ground truth: Transkribus's own guidance recommends transcribing at least 25 pages — between 5,000 and 15,000 words — before training a model. For a lone researcher with a single collection, that is a project before the project.

LLM/VLM transcription — running the page image through ChatGPT, Claude, or Gemini — is the thing many people now reach for first, and it is the most dangerous for this specific task. General vision-language models are text-first systems that downsample the high-resolution image the reading actually depends on. They produce fluent, plausible transcriptions that contain fabricated readings not present on the page. The recent benchmarking literature captures the shape of the problem: one 2025 study found multimodal LLMs reaching around 1.7% character error rate on modern English handwriting (the IAM dataset) but degrading sharply on historical scripts — a single corpus, still requiring verification, and not evidence that these models are safe on pre-modern vernacular hands. The failure mode is what matters: garbled OCR announces itself, while a confident fabrication reads as smooth prose and slips past proofreading. Why fluent LLM errors are harder to catch than garbled ones explains why this is the expensive kind of mistake.

The principle that governs the choice

Whatever method you use, one principle should govern it: transcribe what is on the page, not what the machine thinks should be there. Preserve the archaic spelling, the strikethrough, the marginal addition, the abbreviation mark. The moment a tool silently "corrects" yo[u]r into your or resolves a macron without telling you, you have lost the ability to reconstruct the source — and you may not notice the loss until a reviewer does. How to verify transcription accuracy sets out a working method for checking a machine draft against the image, prioritising the high-stakes tokens: names, dates, numbers.

Where a purpose-built model fits

This is the stage where a specialist tool earns its place. Leo is a specialist HTR platform built around ATR-1, a zero-shot transcription model for Latin-script material — which is to say, whatever the language on the page, so long as it is written in the Latin alphabet. English wills, French notarial records, Dutch registers, German parish books, Italian and Spanish cursive: the constraint is the writing system, not the language. Its flagship public corpus, ExLatinis, happens to be Latin-language — but that is one corpus among many, not the boundary of what the model reads.

Two things make it fit this stage specifically. First, it is zero-shot: there is no 25-page training step before you can transcribe an idiosyncratic hand, which removes the "project before the project" barrier that keeps individual researchers off specialist HTR. Second, the model is trained to preserve source integrity — it transcribes strikethroughs, marginal additions, expansions, and archaic orthography rather than smoothing them into modern prose. On a randomised 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, at ATR-1's release, it recorded roughly a 5% character error rate — 61% fewer errors than the next-best model tested, with Transkribus's Text Titan I around 13% and the general LLMs substantially higher (full benchmark data here). That benchmark is a single manuscript corpus in one language; treat it as a strong signal, not a universal promise, and verify your own material.

Whichever tool you land on, the machine produces a draft, not a finished transcription. The verification step is not optional. An honest comparison of the transcription tool classes is worth reading if you are still choosing.

Stage four: analyse

Once text is machine-readable, the standard computational methods become available: full-text search across a corpus, named-entity recognition, topic modelling, stylometry, distant reading. These are the techniques that let you ask questions of a body of sources too large to read closely.

The constraint is blunt and worth stating: analysis quality is bounded by transcription quality. Above roughly 10% character error rate, downstream tasks degrade materially — named-entity recognition in particular suffers, because a misread surname is simply a different name. This is why the transcription stage is not a formality to rush through. A corpus transcribed at 5% CER supports analysis that a corpus at 20% CER does not.

Analysis is also where interpretation begins, and it should stay separate from the base transcription. If you translate, modernise, summarise, or extract entities, keep those outputs distinct from the verbatim reading of the page. The source record and your interpretation of it are different objects, and conflating them is how errors of reading become errors of argument. The wider question of what AI can and cannot reliably do in historical research is worth thinking through before you lean on any automated analysis.

Stage five: cite

A transcription that cannot be traced back to its source is not yet a citable source. The final stage is provenance: binding the text to the exact image, shelfmark, and collection it came from, so another scholar can verify your reading.

At the scholarly end, this means encoding in TEI P5 — whose Chapter 12 handles manuscript description — and, at institutional scale, linking images to text through IIIF and packaging with METS/ALTO/PAGE. Full TEI encoding is common in digital-edition projects and uncommon among individual working historians, most of whom use lighter conventions. The layered guide to archival metadata standards helps you decide how much encoding your project actually warrants rather than defaulting to the heaviest schema.

Even without full TEI, the citable minimum is fixed: repository, collection, shelfmark, folio or page, and a reference to the source image, formatted per Chicago's archival conventions or MHRA. If you recorded this metadata back at the organise stage, the citation is already half-written. If you did not, you are now reconstructing it — which is precisely why the organise stage is not optional.

The through-line

Five stages, one discipline. The reason to think of research as a workflow rather than a sequence of unrelated tasks is that the stages constrain one another in one direction: capture bounds transcription, transcription bounds analysis, and the metadata you record early is what makes the source citable at the end. Skip a stage and it does not disappear — it resurfaces later as reconstructed provenance, an un-searchable folder, or an analysis built on a reading no one can verify.

The single point where the discipline is most easily lost is fidelity to the page. Machines now read old hands well enough to change the economics of research, and that is genuinely useful. But the historian's job at every stage is the same as it has always been: to know what is actually written, to preserve it as written, and to be able to prove where it came from. The tools change. The obligation to the source does not.

Frequently Asked Questions

What are the stages of a historical research workflow?

A historical research workflow has five stages that turn a manuscript in an archive into a source you can cite with confidence: capture, organise, transcribe, analyse, and cite. Each stage constrains the next. A badly lit photograph limits transcription; a transcription without provenance cannot be cited; and analysis is only ever as good as the text beneath it. The workflow is the same whether you are reading Elizabethan wills, French notarial minutes, or German parish registers — what changes is the hand and the language, not the discipline. The stage that stalls most individual researchers is transcription, where a difficult hand can hold up months of work.

Why is general-purpose OCR the wrong tool for handwritten manuscripts?

General-purpose OCR — Tesseract, ABBYY FineReader, Google Cloud Vision, Amazon Textract — is engineered for clean modern print and is very good at it, but it is the wrong instrument for manuscript hands. ABBYY's engine explicitly limits its handwriting support to handprinted, non-cursive text, and Tesseract's documentation says the same. On secretary hand or Kurrent, glyph-classification approaches fail because connected letterforms merge into shapes the template was never trained on. Specialist handwritten text recognition (HTR) treats a line of text as a sequence rather than a bag of glyphs, which is why it generalises across a writer's variation and reads old hands where OCR cannot.

Can I use ChatGPT or Claude to transcribe historical documents?

You can run a page image through ChatGPT, Claude, or Gemini, but for historical manuscripts it is the most dangerous method for this specific task. General vision-language models are text-first systems that downsample the high-resolution image the reading depends on, and they produce fluent, plausible transcriptions that contain fabricated readings not present on the page. The failure mode is what matters: garbled OCR announces itself, while a confident fabrication reads as smooth prose and slips past proofreading. One 2025 study found these models reaching around 1.7% character error rate on modern English handwriting but degrading sharply on historical scripts — still requiring verification.

Do I need to train a model before transcribing an unusual hand?

Not with a zero-shot tool. Traditionally, usable accuracy on an idiosyncratic hand has required training or fine-tuning a model on ground truth — Transkribus's own guidance recommends transcribing at least 25 pages, between 5,000 and 15,000 words, before training a model. For a lone researcher with a single collection, that is a project before the project. Leo's ATR-1 is zero-shot for Latin-script material, so there is no 25-page training step before you can transcribe an idiosyncratic hand. The constraint is the writing system, not the language: English wills, French notarial records, Dutch registers, German parish books, and Italian or Spanish cursive all fall within it.

How does transcription accuracy affect what analysis I can do?

Analysis quality is bounded by transcription quality. Above roughly 10% character error rate, downstream computational tasks degrade materially — named-entity recognition in particular suffers, because a misread surname is simply a different name. A corpus transcribed at 5% character error rate supports full-text search, topic modelling, stylometry, and distant reading that a corpus at 20% does not. This is why transcription is not a formality to rush through. Keep interpretation separate from the base transcription too: if you translate, modernise, or extract entities, hold those outputs distinct from the verbatim reading, because conflating them is how errors of reading become errors of argument.

© 2026 Leo Technologies Limited. All rights reserved