Leo vs eScriptorium: A No-Training eScriptorium Alternative for Reading Latin-Script Manuscripts
Compares Leo to eScriptorium/Kraken for transcribing Latin-script manuscripts, arguing eScriptorium requires finding or training a recognition model first while Leo offers ready zero-shot transcription, with honest tradeoffs on model coverage, infrastructure cost, and Leo's own limitations.
Leo Team
July 26, 2026

This is a practical comparison for researchers weighing eScriptorium against Leo — and a case for when a ready, no-training eScriptorium alternative is the better path. eScriptorium is a serious open-source tool, but it ships no recognition model, so reading a page means finding or building one first. If your material is vernacular Latin-script handwriting from the last few centuries, that first step is often the whole bottleneck.
eScriptorium is a capable, open-source platform for transcribing historical manuscripts — but it ships no general-purpose recognition model, so out of the box it transcribes nothing. To get text off a page you must first obtain a trained Kraken model that matches your language and hand, or build one from ground truth yourself. If your material is medieval Latin or French, a strong public model already exists. If it is English wills, Dutch registers, or German Kurrent, you are largely on your own. Leo is the alternative for researchers who need to read Latin-script hands now, without training a model first: a ready, zero-shot transcription engine wrapped in a document workspace.
That is the whole decision. The rest of this article works through it honestly, because "no training" is a claim worth testing carefully, and eScriptorium is a serious tool that deserves an accurate description rather than a straw man.
What eScriptorium Actually Is
The most common misconception about eScriptorium is that it is a transcription engine. It is not. eScriptorium is a web platform for manual and automated segmentation and text recognition — a Django application that acts as a graphical interface and orchestration layer over a separate recognition engine called Kraken. Kraken does the reading. eScriptorium is where you manage images, correct output, and organize training data.
This matters because of what Kraken is: a turn-key OCR/HTR library "optimized for historical and non-Latin script material," and, crucially, script- and language-agnostic by design. That agnosticism is a genuine strength. Kraken will learn Greek, Hebrew, Arabic, or a Carolingian minuscule with equal willingness. But agnosticism has a cost: the system knows nothing until you teach it. eScriptorium ships Kraken's default segmentation model, which finds text lines on a page, but no general-purpose recognition model comes with the install. Recognition — turning those lines into characters — requires a model you supply.
So the real eScriptorium workflow has two forks. Either you find a public model that matches your material, load it, pair it with a segmentation model, and validate it on a small in-domain sample — or you train your own from scratch.
The Two Paths, and Where Each One Stalls
Path One: Find a Public Model
The open ecosystem for Kraken models — chiefly HTR-United and Zenodo deposits — is real and growing. If your manuscripts fall inside its coverage, you can reach usable transcription without training anything. That is the "no-training" case, and it is worth taking seriously.
The problem is coverage. The most mature and most-cited open deposit is CREMMA-Medieval, covering medieval Latin (11th–16th century) and Old French (13th–15th century) across 2,828 folders and 86,832 transcription lines. French and medieval Latin, in other words, are well served. Everything else is uneven. English, Dutch, Italian, and Spanish coverage is thinner and patchier. German Kurrent and Sütterlin are, as one honest census puts it, essentially do-it-yourself except through vendor-private models.
Note carefully that this is coverage by language, not merely by script — which is the point where the "no-training" promise gets specific. Kraken can read the Latin alphabet in principle. Whether your Latin-script material — a seventeenth-century English will in secretary hand, a nineteenth-century Dutch parish register — has a public model behind it is a separate question, and for most vernacular material written in the past 500 years, the honest answer is "not really, not yet." For a fuller treatment of why script, language, and hand are three different constraints, see our guide to which languages and scripts HTR can read.
And even when a public model exists, deployment is not one click. The model must be loaded, matched to a segmentation model, validated against your own sample, and — very often — fine-tuned before it performs on your specific hand. A public model is a head start, not a finished job.
Path Two: Train Your Own
If no public model fits, you build one. This is where eScriptorium's design shows its real character: it is, at heart, an excellent environment for producing ground truth — the per-line pairs of image and correct transcription (in ALTO or PAGE XML) that a model learns from.
The volumes involved are well documented. Kraken's own training guidance notes that most western texts run 25–40 lines per page, so "upward of 30 pages" must be preprocessed and transcribed to start — roughly 750 to 1,200 lines. The current training tutorial works through an example built on 876 lines. Accuracy seldom improves after 50 epochs, which take between 8 and 24 hours on a normal desktop; characters appearing fewer than ten times in your ground truth will most likely not be recognized well. Fine-tuning the segmentation model for a book-specific layout wants 30–50 labelled pages, with stronger results around 1,000.
None of this is a criticism of eScriptorium. It is simply what training a model is. For a research group with a large, homogeneous corpus in a single hand — the exact case eScriptorium and Kraken were built to serve — that investment amortizes well across thousands of pages. The trouble is that a great deal of archival research does not look like that. It looks like forty pages here in one clerk's hand, sixty there in another, a bundle of letters in a third. Every distinct hand potentially restarts the ground-truth clock, and the payoff per model shrinks.
The Cost That "Free" Hides
eScriptorium is free of licence fee, and open-source software genuinely lowers a real barrier. But "free" describes the licence, not the total cost of ownership. The recommended deployment is via Docker containers; running it in earnest means Linux, 8–16 GB of RAM, and — for training at any speed — an NVIDIA GPU. Installation is not always smooth: a long-standing GitLab issue documents the migration pain from the obsolete `nvidia-docker2` to the current container toolkit for GPU support. Add IT time to stand it up and maintain it, plus the researcher-weeks spent building ground truth, and the true cost is measured in infrastructure and hours, not dollars.
For a well-resourced digital-humanities lab with technical staff, this is entirely manageable and often the right choice — the control and reproducibility of a self-hosted, open pipeline are real scholarly goods. For an individual historian or a small team whose bottleneck is reading the documents, not building infrastructure, the calculus is different.
Where Leo Fits: A Ready Model, No Training Step
This is the specific gap Leo is built to close. Where eScriptorium requires you to obtain or train a recognition model before you can transcribe anything, Leo's engine, ATR-1, is a specialized, zero-shot transcription model that reads Latin-script material out of the box — manuscript hands including secretary, cursive, and court hands, as well as printed matter. No ground truth to assemble, no segmentation model to pair, no GPU to provision, no epochs to wait through. You upload page images — including straight from a phone camera in the reading room — and transcribe.
The scope is defined by the writing system, not the language: any language written in the Latin alphabet is in scope, with performance strongest in English and strong across French, German, Spanish, Italian, Dutch, and Latin among others. This directly addresses the coverage gap above — the English wills, Dutch registers, and other vernacular material where the open-model ecosystem is thin, and where the eScriptorium path most often means training from scratch. Transcription is not translation: Leo transcribes what is on the page; rendering it into modern English is a separate, one-click Translate operation that writes to a new tab, leaving the base transcription untouched.
On head-to-head accuracy, the one cleared, published comparison is worth reading in full. On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, at ATR-1's release, Leo scored roughly 5% character error rate — 61% fewer errors than the next-best model (Transkribus/Text Titan I ~13%; Claude Opus ~23.3%; Gemini 2.5 Pro ~24.8%; GPT-4.1 ~56.7%). One caveat travels with that figure and should stay attached to it: it is a single corpus, one language, one period, benchmarked at a point in time — not a universal accuracy guarantee. As with any machine transcription, the output is a first-pass draft to verify against the image, not a finished edition. But it is a like-for-like measurement on exactly the kind of vernacular early-modern hand where the eScriptorium path is hardest.
There is a second difference, and for daily work it may matter more than the accuracy line. eScriptorium is a recognition-and-annotation platform; organizing a collection, describing it, and analyzing it happen elsewhere. Leo folds the surrounding workflow into one place — upload, organize into folders, attach structured metadata, search across everything with fuzzy matching, edit in a rich-text editor that preserves strikethroughs, insertions, and marginalia, and export to TEI XML, Word, PDF, or HTML. On top of a faithful base transcription, a separate AI layer can summarize, classify, extract named entities, or translate — each result landing in its own tab, so interpretation is layered on the source rather than baked into it.
What Leo does not do is worth stating plainly, because it marks the honest boundary of this comparison. Leo offers no user model-training or fine-tuning: there is no path to teach it your specific idiosyncratic hand the way Kraken lets you. It exports TEI, not ALTO or PAGE XML, so it is not a drop-in producer of the line-level ground-truth artifacts an eScriptorium pipeline consumes. And its sharpest known weak spot is pages that mix dominant printed structure with dense handwriting — pre-printed ledger and deed forms — where it can favor the printed headers over the manuscript entries. If your project genuinely requires a custom-trained, reproducible open-source model over a homogeneous corpus, eScriptorium is the right tool, and Leo is not pretending otherwise.
How to Choose
Reduced to essentials, the decision turns on three questions.
Does a Public Kraken Model Already Cover Your Exact Material?
If you work in medieval Latin or Old French, quite possibly yes — and the "no-training" case for eScriptorium is real. Load the model, validate it on a sample, and you may need no custom training at all. If you work in English, German, Dutch, Spanish, or Italian vernacular hands from the last few centuries, probably not, and you are looking at building ground truth.
Do You Have the Infrastructure and the Corpus to Justify Training?
A technical team with a large, homogeneous body of material in a consistent hand will get excellent, reproducible results from a self-hosted eScriptorium/Kraken pipeline, and full control over it. An individual researcher moving across many hands, whose bottleneck is reading rather than engineering, will spend that same effort on setup instead of scholarship.
Where Is Your Time Better Spent?
This is the real question underneath the other two. Every hour building ground truth or wrestling a Docker GPU install is an hour not spent reading, analyzing, and writing. If model-building is itself part of your project — a methods contribution, a shared resource for a field — that time is well spent. If transcription is merely the wall between you and your sources, a ready zero-shot engine removes the wall without asking you to become a machine-learning practitioner first.
Neither answer is universally right. The open-source, trainable path and the ready-model path solve different problems, and a researcher who understands the difference will choose well. For a broader map of how these tools sit against general OCR, large language models, and collection managers, our comparison of HTR software and alternatives lays out the full picture.
Whatever engine reads your first draft, the transcription is not finished when the machine stops. It is finished when you have checked it against the image — attending first to the names, dates, and figures where a single wrong character does the most damage — and made the editorial decisions no model can make for you. The tool that gets you to a verifiable draft fastest, without demanding weeks of setup you did not come to the archive to do, is the one that returns the most hours to the work that only a historian can do: reading these sources closely, and understanding what they say.
Frequently Asked Questions
What is a good eScriptorium alternative that doesn't require training a model?
Leo is an eScriptorium alternative built for researchers who need to read Latin-script hands now, without training a model first. Where eScriptorium ships no general-purpose recognition model — meaning you must find a public Kraken model or build one from ground truth before you can transcribe anything — Leo's engine, ATR-1, is a zero-shot transcription model that reads Latin-script material out of the box, including secretary, cursive, and court hands as well as printed matter. You upload page images, even straight from a phone camera, and transcribe. No ground truth, no segmentation model to pair, no GPU, no training epochs.
Does eScriptorium come with a transcription model built in?
No. eScriptorium ships no general-purpose recognition model, so out of the box it transcribes nothing. It is a web platform for manual and automated segmentation and text recognition — a graphical interface and orchestration layer over a separate engine called Kraken, which does the actual reading. eScriptorium does include Kraken's default segmentation model, which finds text lines on a page, but turning those lines into characters requires a recognition model you supply. That means either loading a public Kraken model that matches your language and hand, or training your own from ground truth in ALTO or PAGE XML.
How much ground truth do you need to train a Kraken model in eScriptorium?
Kraken's own training guidance suggests starting with upward of 30 pages of transcribed material — roughly 750 to 1,200 lines, since most western texts run 25–40 lines per page. The current training tutorial works through an example built on 876 lines. Accuracy seldom improves after 50 epochs, which take between 8 and 24 hours on a normal desktop, and characters appearing fewer than ten times in your ground truth will most likely not be recognized well. Fine-tuning the segmentation model for a book-specific layout wants 30–50 labelled pages, with stronger results around 1,000.
Which languages does the open Kraken model ecosystem cover well?
The open ecosystem for Kraken models covers medieval Latin and Old French well, but is uneven everywhere else. The most mature and most-cited deposit, CREMMA-Medieval, covers medieval Latin from the 11th–16th century and Old French from the 13th–15th century across tens of thousands of transcription lines. English, Dutch, Italian, and Spanish coverage is thinner and patchier, while German Kurrent and Sütterlin are essentially do-it-yourself except through vendor-private models. This is coverage by language, not just by script: Kraken can read the Latin alphabet in principle, but whether your specific vernacular material has a public model behind it is a separate question.
Is eScriptorium really free to use?
eScriptorium is free of licence fee and open-source, but "free" describes the licence, not the total cost of ownership. The recommended deployment is via Docker containers, and running it in earnest means Linux, 8–16 GB of RAM, and — for training at any speed — an NVIDIA GPU. Installation is not always smooth; GPU support has documented migration pain around the container toolkit. Add IT time to stand it up and maintain it, plus the researcher-weeks spent building ground truth, and the true cost is measured in infrastructure and hours, not dollars. For a well-resourced lab this is manageable; for an individual researcher it may not be.