Does AI Mean Historians Still Need Paleography? The Skill That Now Verifies the Machine
How AI transcription changes paleography for historians: it can speed first-pass reading, but human verification, editorial judgment, and document-context expertise remain essential.
Leo Team
July 20, 2026

This article answers a question every working historian now faces: do historians still need paleography when AI can transcribe a manuscript in seconds? The short answer is yes — but the skill has moved, from first-pass reading to verification. Understanding where it moved is what keeps your transcriptions defensible.
No. On historical handwriting, general-purpose AI still fails badly enough that a scholar with paleographic competence remains the point at which a transcription becomes trustworthy. Machine transcription has genuinely changed the shape of the work — it removes much of the line-by-line decoding burden — but it has not removed the skill. It has relocated it, from first-pass reading to verification, editorial judgment, and knowing when the machine is confidently wrong. The published evidence is consistent on this point, and the direction of travel is not "paleography is obsolete" but "paleography is now the thing that makes AI safe to use."
That is the short answer. The longer one is worth having, because why the skill survives tells you which parts of it now matter most.
What paleography actually is — and why "reading old writing" undersells it
It helps to be precise about the discipline before asking whether a machine can replace it. Paleography is the study of historical handwriting: not only deciphering scripts, but dating and localizing hands, identifying scribal practice, expanding abbreviations and brevigraphs, and reading a text as a material artifact. It sits alongside two neighboring disciplines that the "just read the letters" framing ignores entirely. Diplomatics — founded by Mabillon's De re diplomatica in 1681 — is the critical analysis of how a document was produced: its formulae, protocols, and conventions. Codicology is the study of the physical codex as object.
None of these three is a decoding task. A machine can be trained to turn marks on a page into characters. It cannot, on its own, tell you that a formula is anomalous for the chancery that supposedly issued it, that a hand is a generation too late for its claimed date, or that an abbreviation should be expanded one way in a legal instrument and another in a personal letter. Those are judgments, and they are the part of the discipline that survives any conceivable improvement in recognition accuracy.
This matters for the AI question because the loudest version of the claim — that AI has made paleography obsolete — quietly redefines paleography as decoding alone. Once you hold the fuller definition in view, the claim collapses before you even look at the error rates. But the error rates are worth looking at anyway, because they settle the narrower question too.
What the evidence actually shows
The narrow claim is testable: can a general-purpose AI model read a historical hand well enough to skip the paleographer? The published benchmarks say no, and they say it clearly.
The most rigorous test to date is Crosilla, Klic and Colavizza's 2025 benchmark of large language models on handwritten text recognition. On clean modern handwriting — the IAM English dataset — GPT-4o reaches roughly 1.4–1.75% character error rate, genuinely competitive with specialist systems. This is the result that fuels the "AI can read anything now" impression, and on modern material it is broadly fair. But on the ICDAR2017 set of historical German handwriting, every general-purpose model tested collapsed: character error rates between 41% and 86%, with even the strongest, Claude 3.5 Sonnet, at 41.19%. A specialist supervised model on the same historical test set reached 7.07% CER. The gap is roughly an order of magnitude, and — this is the important part — it is driven by the corpus, not the model. The same GPT-4o that scores ~1.7% on modern English falls to around 7% on historical English. Move the material back a few centuries and the machine's confidence outruns its competence.
Character error rate, for readers newer to the metric, is edit-distance-based: a CER of 5% means roughly five wrong characters per hundred. At 40–80%, you are not reading a manuscript. You are reading the model's guess at what a manuscript like that might plausibly say — which is a different and more dangerous thing.
That danger is the second finding worth carrying. The University of Virginia Library's February 2026 staff experiment tested generalist AI tools across a 600-year span of material — an 1835 letter, a nineteenth-century police guard book, an early-modern Spanish document, and more. The dominant failure mode was not garble. It was hallucination: fluent, plausible, confidently wrong readings. Crosilla and colleagues found the same models "do not possess a significant capability for self-correction." This is why the safety argument matters more than the accuracy argument. A garbled OCR line announces its own failure; a fabricated but grammatical sentence does not. The error that reads smoothly is the error that survives into your footnotes. We have written separately about why fluent LLM transcription errors are so much harder to catch than the garbled kind, and why ChatGPT in particular struggles with historical handwriting — but the short version is that plausibility is the trap.
Two honest caveats belong here, because the evidence base is thinner than the confident headlines on either side suggest. The LLM-hallucination finding rests substantially on one rigorous benchmark plus a small staff experiment; it needs independent replication before anyone treats it as settled law. And there is no ImageNet-equivalent standardized benchmark for historical handwriting, so cross-study CER figures are measured against idiosyncratic held-out sets and cannot be compared head-to-head with any precision. What the evidence supports is a direction, not a decimal place: general AI is unreliable on historical hands, and its errors are the dangerous kind.
The claim that quietly proves the point
There is a detail in the specialist-HTR literature that settles the "raw output is publication-ready" version of the question almost by itself. Even a well-tuned specialist model on a well-matched hand lands somewhere in the 1.8–7% CER range — that is, several errors per hundred characters, every hundred characters, across a whole document.
The clearest evidence that this is not good enough on its own is that people build correction pipelines on top of it. Transcription Pearl (Humphries et al., November 2024) pipes Transkribus output through Claude Sonnet 3.5 to reach a modified CER around 1.8%. The existence of that pipeline is the argument: if specialist first-pass output were publication-ready, no one would engineer a second stage to clean it up. And note what the second stage is doing — it is contextual correction, exactly the kind of judgment a paleographer applies, now partly automated but still requiring a human to check that the "correction" did not smooth a genuine reading into a plausible false one.
There is a deeper dependency underneath all of this. Every published HTR pipeline — specialist or hybrid — was trained or fine-tuned on ground truth: human-verified transcriptions produced by people who can already read the hand. The data the model depends on is the product of the very skill the model is said to make redundant. The discipline is not upstream of the technology by accident. It is upstream by necessity.
What actually changes: from decoding to verification
So paleography survives — but it does not survive unchanged, and pretending otherwise would be as misleading as the obsolescence claim. The honest description is that machine transcription reorganizes the labor. The bottleneck used to be first-pass reading: hours spent decoding secretary hand letter by letter before analysis could even begin. A good HTR first pass compresses that. What it does not compress — what it arguably increases — is verification.
The skills that now matter most are the ones a machine cannot supply:
- Adjudicating uncertain readings. The machine gives you a character string; you decide whether it is right, and you can only decide if you can read the hand yourself. Verification is not a lighter form of paleography. It is paleography applied to someone else's transcription.
- Editorial judgment. Whether to produce a diplomatic transcription preserving the source exactly, a semi-diplomatic one expanding abbreviations in brackets, or a normalized one modernizing spelling, is an editorial decision — made by the scholar, never by the model. So is the decision about whether to expand or preserve a given abbreviation, which turns on genre, audience, and argument.
- Knowing the hand and the tradition. The abbreviation systems of English secretary hand, French and Italian notarial cursive, German Kurrent, and Spanish procesal each carry their own conventions. Recognizing when a machine has misread within one of those traditions requires knowing the tradition. Our working guide to early modern paleography covers the hands and letterforms where this matters most.
- Diplomatic and codicological reading. Everything about production, dating, and material form that recognition simply does not touch.
This is the substance behind the broader principle that runs through careful practice with these tools: keeping a human in the loop is not a compliance gesture but the mechanism by which fluent output becomes defensible scholarship. Whether HTR lowers the entry barrier for working with old documents or merely shifts the effort from decoding to validation is itself an open empirical question — the UVA experiment is suggestive, not conclusive. But for the working historian the practical implication is the same either way: you still need to be able to read the page.
Where a specialist tool fits — and where it does not
If verification is the surviving skill, the right tool is one that gives you the least-corrupted starting point and keeps your judgment in charge of the record. This is a real distinction between tool families, and it is worth being precise about it rather than lumping "AI" into one bucket.
A general LLM is the wrong instrument here for the reason the evidence above makes plain: it downsamples the high-resolution image the reading actually depends on, it has no interface built for the transcription workflow, and it fails toward fluent fabrication — the error you are least able to catch. A specialist HTR model errs differently. When it is wrong, it tends to be wrong at the level of characters and words: recoverable mistakes that a paleographer's eye catches, not invented sentences that read like the real thing.
This is the stage where a purpose-built model earns its place. Leo's transcription model, ATR-1, is built for exactly this: Latin-script manuscripts and printed matter across English, French, German, Dutch, Spanish, Italian, Latin and the other languages written in that alphabet — the historical hands where general OCR and general LLMs both fail. Its design commitment is source integrity: it transcribes what is on the page rather than smoothing it into modern, plausible prose, preserving strikethroughs, insertions, marginal notes, and archaic orthography rather than silently "correcting" them. On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, ATR-1 scored roughly 5% character error rate at release — about 61% fewer errors than the next-best model tested, with Transkribus's Text Titan I at ~13% and general LLMs far higher (full benchmark data here). Unlike the specialist HTR tools that require you to build and train a model on your own corpus first, it runs zero-shot, out of the box — which matters precisely because the ground-truth-building step is itself skilled paleographic labor.
But notice what that ~5% figure means, and what this section is not claiming. Five errors per hundred characters is a strong first pass, not a finished transcription. The tool's job is to hand you the cleanest possible draft and then keep the base transcription faithful to the page while you verify it — corrections you make feed back into later versions of the model, but the judgment stays yours. It does not adjudicate a doubtful reading, choose your transcription convention, or tell you the diplomatic formula is wrong for the archive. Those remain the paleographer's work. A specialist tool changes where you spend your hours; it does not change whose competence the final text depends on.
The honest conclusion
Does AI mean historians no longer need paleography? Only if you believe paleography was ever just decoding — and it never was. What has actually happened is narrower and more interesting. The machine has taken over much of the first-pass reading, and in doing so it has raised the value of everything the machine cannot do: verifying an uncertain reading against the hand, choosing how to represent the text, recognizing when a fluent transcription is confidently wrong, and reading the document as the material, produced object it is.
The scholar who treats AI output as finished is the scholar whose errors will be the hardest to detect and the most damaging to correct. The scholar who can read the page will use these tools well, catch what they get wrong, and produce work that holds up. Paleography is not the skill AI made redundant. It is the skill that tells you when to trust the machine and when not to — and that skill is worth more now, not less.
Frequently Asked Questions
Do historians still need paleography now that AI can transcribe manuscripts?
Yes. On historical handwriting, general-purpose AI still fails badly enough that a scholar with paleographic competence remains the point at which a transcription becomes trustworthy. Machine transcription has changed the shape of the work — it removes much of the line-by-line decoding burden — but it has relocated the skill rather than removed it, from first-pass reading to verification, editorial judgment, and knowing when the machine is confidently wrong. Paleography is now the thing that makes AI safe to use: it tells you when to trust the machine and when not to.
What is paleography, and is it just reading old handwriting?
Paleography is the study of historical handwriting — not only deciphering scripts, but dating and localizing hands, identifying scribal practice, expanding abbreviations, and reading a text as a material artifact. Calling it "reading old writing" undersells it. It sits alongside diplomatics, the critical analysis of how a document was produced through its formulae and conventions, and codicology, the study of the physical codex as object. None of these three is a decoding task. A machine can turn marks into characters, but it cannot tell you a formula is anomalous for its chancery or that a hand is a generation too late for its claimed date.
How accurate is AI at reading historical handwriting?
Poorly, on historical hands. General-purpose models tested on a set of historical German handwriting collapsed to character error rates between 41% and 86%, while a specialist supervised model reached about 7% on the same material — roughly an order of magnitude better. The gap is driven by the corpus, not the model: GPT-4o scores around 1.4–1.75% on clean modern English handwriting but falls to about 7% on historical English. Move the material back a few centuries and the machine's confidence outruns its competence. Character error rate is edit-distance-based, so a 5% rate means roughly five wrong characters per hundred.
Why are AI transcription errors on manuscripts so dangerous?
Because they read smoothly. The dominant failure mode of generalist AI on historical documents is not garble but hallucination: fluent, plausible, confidently wrong readings. A garbled OCR line announces its own failure; a fabricated but grammatical sentence does not — it survives quietly into your footnotes. Researchers have also found these models do not possess a significant capability for self-correction. This is why the safety argument matters more than the accuracy argument. The error that reads plausibly is the error you are least able to catch, which is exactly why a scholar who can read the hand must verify the output.
What skills does a historian need if AI does the first-pass reading?
Verification, above all. Machine transcription reorganizes the labor: it compresses first-pass decoding but increases the work of checking. The skills that matter most are ones a machine cannot supply — adjudicating uncertain readings against the hand, which you can only do if you can read the hand yourself; editorial judgment about whether to produce a diplomatic, semi-diplomatic, or normalized transcription; knowing the abbreviation systems of specific traditions like secretary hand, notarial cursive, Kurrent, or procesal; and reading the document diplomatically and codicologically. Verification is not a lighter form of paleography — it is paleography applied to someone else's transcription.