AI and the Humanities: How Computation Supports Historical Scholarship Without Replacing It
AI in the humanities as a controlled scholarly tool for transcription, pattern-finding, and translation, with historians retaining verification, interpretation, and source criticism.
Leo Team
July 20, 2026

This article examines how AI belongs in the humanities — as a specialized instrument under scholarly control, not as an author. For historians working with manuscripts and large corpora, the distinction between support and replacement is the difference between accelerating the work and manufacturing plausible error at scale.
AI in the humanities is best understood as an instrument, not an author. In practice this means the machine handles the mechanical passes — turning manuscript images into searchable text, surfacing thematic patterns across a large corpus, drafting a translation — while the historian retains judgment, verification, and interpretation. The consensus across professional bodies, most recently the American Historical Association's Guiding Principles for AI in History Education (August 2025), is that AI operates under explicit scholarly control, with disclosed methods and verifiable sources. Where that discipline holds, computation accelerates the work. Where it slips, it manufactures plausible error at scale.
That distinction — support versus replacement — is not new anxiety dressed in new technology. It is the oldest question in the humanities-computing tradition, and the history of the field is largely a record of getting the answer right.
A lineage of computation as instrument
The through-line runs back further than most current debate acknowledges. In 1949, the Jesuit scholar Roberto Busa began what became the Index Thomisticus, a concordance of the complete works of Thomas Aquinas. Over thirty-four years it processed roughly 10.6 million words, published across fifty-six volumes between 1974 and 1980. Busa did not ask the machine to interpret Aquinas. He asked it to do what a human indexer would do too slowly to be useful: find every instance of a word and its context, so that a scholar could then read.
The same logic organizes the field's later turns. Franco Moretti's distant reading proposed computational, aggregate analysis of large literary corpora as a complement to close reading — not a substitute for it. Stylometry, built on function-word frequencies, matured into a defensible method through Burrows's Delta (2002), with the Federalist Papers as its canonical case. In each case the pattern is identical: computation proposes; the scholar disposes. The instrument extends reach. It does not confer understanding.
This is the frame worth holding as generative AI enters the same lineage. The tools are more capable and, in one specific way, more dangerous. But the governing principle has not changed. It belongs to the broader conversation about AI in the humanities and cultural heritage, and it is the standard against which any new tool should be measured.
Where AI genuinely supports the work
It helps to be specific about which tasks AI now handles well, which it handles adequately under supervision, and which it should not be trusted with at all. The mechanism families below are all in routine scholarly use. They are not equally mature.
Reading the page: HTR and OCR
The most mature application is text recognition. Handwritten text recognition (HTR) uses machine learning to transcribe manuscript hands; optical character recognition (OCR) does the equivalent for printed matter. On trained models, HTR platforms such as Transkribus report 90–98% character accuracy for English, German, French, Dutch, and Spanish hands — turning collections that were previously legible only to specialists into searchable text.
The qualifier "on trained models" carries weight. Accuracy on cursive or historical hands without training falls below 50%, and the gap between a headline figure and a real manuscript is where most disappointment lives. The Folger Shakespeare Library's own experiments with early-modern English material recorded a character error rate of roughly 21.25% — about one error in five symbols — on a difficult hand. The lesson is not that HTR fails. It is that accuracy depends on how closely the model matches the material in front of you. The mechanics of that dependence are worth understanding in detail, and we cover them in a look inside the HTR pipeline and in the practical difference between HTR and OCR.
Historical print deserves particular mention, because the common assumption that "OCR handles old books fine" is wrong. General-purpose OCR is engineered for clean modern type. Early-modern typography defeats it systematically: the long s read as an f, ligatures split or dropped, blackletter and Fraktur founts it has never seen, uneven inking and show-through read as character evidence. Character error rates that sit below 2% on modern print rise to 15–40% on sixteenth- to eighteenth-century editions without specialised handling — which is precisely why projects like OCR4all and the Early Modern OCR Project exist. The point that generalises: recognition quality is a function of the match between model and material, not of a marketing language count. We work through what that means for scripts and languages in a guide to which languages and scripts HTR can actually read.
Finding patterns: topic modelling and NER
Above the level of the individual page, two families of tools help scholars work at corpus scale.
Topic modelling — LDA, MALLET, and their relatives — surfaces thematic clusters across a large body of text. It remains in routine use, with one standing caveat that its practitioners repeat because it matters: the outputs are distributional, not causal. As Benjamin Schmidt argued in "Words Alone", a topic model reflects the statistics of vocabulary, not historical causation, and must be read back against the documents themselves. Treated as a finding aid, it is valuable. Treated as an answer, it misleads.
Named entity recognition — tagging persons, places, and organizations — is feasible but not yet mature on historical material. Benchmarks on eighteenth-century text report F1 scores below 70% even under lenient evaluation. That is useful for generating candidate indexes a scholar then checks. It is not reliable enough to publish unverified. This layered approach — machine draft, human verification, interpretation kept separate from the source — is the working pattern behind sound historical document analysis.
Scaling human review: crowdsourced transcription
The dominant model for large transcription projects remains hybrid: an algorithmic first pass, then volunteer or specialist verification. Transcribe Bentham, Zooniverse, and FromThePage all embed the same human-in-the-loop principle the professional bodies now treat as the ethical baseline. The machine proposes; the human verifies and interprets. Nothing in the current generation of tools has retired that division of labour. It has only made the first pass faster.
The failure mode that matters most
There is one application where the "support, not replace" line is not a preference but a safety requirement, and it is the one scholars most often reach for first: asking a general-purpose LLM to transcribe a manuscript or to supply references.
The problem is structural, not a matter of an immature model that will improve next quarter. Large language models are trained to produce plausible, fluent text. On a task that demands fidelity to marks on a page, plausibility is exactly the wrong objective. The model does not fail loudly by producing garbage. It fails quietly by producing something that reads correctly and is wrong. The AHA's principles put it plainly: generative AI "regularly hallucinates content, references, sources, and quotations," and its output must be verified against the original.
The citation evidence is stark. A Deakin University study (November 2025) found ChatGPT fabricated close to 20% of citations outright and introduced errors into roughly 45% of the real references it cited across simulated literature reviews. A separate systematic-review study by Chelli and colleagues reported fabrication rates as high as 28.6%. These are not edge cases. They are the predictable behaviour of a system optimised for fluency rather than truth.
The same dynamic governs transcription. A general model, handed a photograph of a seventeenth-century hand, will confidently return modern, grammatical prose — smoothing archaic orthography, resolving abbreviations it has guessed at, and occasionally inventing a clause that fits the surrounding sense. This is worse than an obvious error, because it is hard to catch. A specialist recognition error tends to land on a character or a word: a misread letter, a dropped mark, the kind of mistake a reader scanning against the image will spot. A hallucination lands on meaning, dressed in plausible fluency. We examine why this happens — image downsampling, fluent confabulation — in a dedicated piece on why ChatGPT struggles to transcribe handwriting.
This is where the choice of instrument stops being neutral. For transcription specifically, the safer path is a model built for the task — one trained to weigh high-resolution visual evidence against context and, critically, trained to transcribe what is on the page rather than to normalise it into modern prose. This is the design principle behind Leo's ATR-1: source integrity as the governing commitment. Strikethroughs, marginal additions, editorial expansions, and archaic spelling survive the transcription rather than being silently corrected away. On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, ATR-1 recorded roughly a 5% character error rate at release — 61% fewer errors than the next-best model tested, with Transkribus's Text Titan I at about 13% and the general LLMs ranging from 23% to 57% (benchmark data here). It reads Latin-script material in whatever language the page carries — English wills, French notarial records, Dutch registers, German parish books, and Latin among them. The figure worth internalising is not the accuracy number but the shape of the error: recoverable characters, not fabricated meaning.
The right role, even here, is support. The model produces a faithful first pass; the historian verifies it against the image and interprets what it says. Any downstream transformation — modernization, translation, entity extraction — is kept as a separate layer, so interpretation is never baked into the base record. That separation is what keeps the work defensible.
The open questions worth keeping in view
Honest treatment of AI in the humanities means naming what is not settled. Several questions remain genuinely open, and a scholar adopting these tools should hold them consciously.
The first is deskilling. Does routine reliance on AI for sourcework erode the close-reading and source-criticism capacities of students and early-career researchers? The AHA flags the concern; no longitudinal study has yet measured it. The second is bias. Training data skews English-dominant, modern-print-dominant, and elite-archive-dominant, and when LLMs are used for synthesis rather than transcription, the direction of that distortion is recognised but its magnitude is unmeasured. The third is reproducibility: non-deterministic models and versioned APIs make it unclear whether a computational result in historical scholarship can be reliably reproduced, and no agreed standard yet exists.
There is also the concern, raised by Schmidt and by Tim Hitchcock, that corpus-scale analysis flattens nuance and risks silencing voices under-represented in digitised collections. And there is the labour question the Royal Historical Society and others continue to deliberate: where the boundary lies between honest citation error and misconduct when a model fabricates a reference, and what automation means for the research assistants and transcriptionists whose tasks are most exposed. None of these has a settled answer. All of them are reasons to keep the human firmly in the loop, not to abandon the tools.
The historian's judgment is the point
The instruments have changed since Busa fed punched cards through an IBM machine to concord Aquinas. The principle has not. Every capable tool in this field — HTR, topic modelling, entity extraction, machine translation — earns its place by extending what a scholar can read, search, and compare, and forfeits it the moment it is asked to supply judgment the scholar has not exercised. The machine's proper output is a faithful draft and a set of candidates. The interpretation, the source-criticism, the decision about what a document means and how far it can be trusted — these remain the historian's work, and the quality of the scholarship still rests on how well that work is done. Used that way, AI does not replace the discipline. It gives the discipline more of what it has always been short of: time to read.
Frequently Asked Questions
How is AI used in the humanities?
AI in the humanities works as an instrument, not an author. It handles mechanical passes — turning manuscript images into searchable text through handwritten text recognition, surfacing thematic patterns across large corpora through topic modelling, drafting translations, and tagging entities — while the historian keeps judgment, verification, and interpretation. This division runs back to Roberto Busa's 1949 concordance of Aquinas and through distant reading and stylometry: computation proposes, the scholar disposes. Professional bodies, including the American Historical Association, hold that AI should operate under explicit scholarly control, with disclosed methods and verifiable sources. Where that discipline holds, computation accelerates the work.
Can AI accurately transcribe historical handwriting?
Trained handwritten text recognition (HTR) models can reach 90–98% character accuracy on English, German, French, Dutch, and Spanish hands, but accuracy depends on how closely the model matches the material. On cursive or historical hands without training, accuracy falls below 50%; the Folger Shakespeare Library recorded roughly a 21% character error rate on a difficult early-modern hand. A model built for the task fares far better: Leo's ATR-1 recorded about a 5% character error rate on a 97-image Folger sample, ahead of Transkribus and general LLMs. Recognition quality is a function of the match between model and material, not a marketing language count.
Why shouldn't I use ChatGPT to transcribe manuscripts or find sources?
General-purpose large language models are trained to produce plausible, fluent text, which is the wrong objective for tasks demanding fidelity to marks on a page. They fail quietly, returning something that reads correctly and is wrong. A Deakin University study found ChatGPT fabricated close to 20% of citations and introduced errors into roughly 45% of real references; a separate study reported fabrication rates as high as 28.6%. Handed a seventeenth-century hand, a general model smooths archaic spelling, resolves abbreviations it has guessed at, and can invent clauses. That hallucination lands on meaning, dressed in fluency, which is far harder to catch than a specialist recognition error.
What is the difference between HTR and OCR?
Handwritten text recognition (HTR) uses machine learning to transcribe manuscript hands, while optical character recognition (OCR) does the equivalent for printed matter. The common assumption that OCR handles old books fine is wrong: general-purpose OCR is engineered for clean modern type, and early-modern typography defeats it systematically — the long s read as an f, ligatures split or dropped, blackletter and Fraktur founts it has never seen, uneven inking and show-through misread as characters. Character error rates that sit below 2% on modern print rise to 15–40% on sixteenth- to eighteenth-century editions without specialised handling, which is why projects like OCR4all exist.
Will AI replace historians?
No. Every capable tool in the field — HTR, topic modelling, entity extraction, machine translation — earns its place by extending what a scholar can read, search, and compare, and forfeits it the moment it is asked to supply judgment the scholar has not exercised. The machine's proper output is a faithful draft and a set of candidates. The interpretation, the source-criticism, and the decision about what a document means and how far it can be trusted remain the historian's work, and the quality of scholarship still rests on how well that work is done. Used this way, AI gives the discipline more time to read.