Crowdsourced Transcription: What Volunteers Do Best, and Where HTR Should Take the First Pass
What crowdsourced transcription volunteers reliably deliver, what programme figures do not prove, and how to decide series by series when HTR should take the first pass.
Leo Team
August 10, 2026

Contents
Crowdsourced transcription is how most heritage, court-record, and data-rescue backlogs are being worked through today — and for good reasons. This article sets out what volunteer programmes reliably deliver, what the published figures do and do not prove, and how to decide, series by series, when handwritten text recognition should produce the first draft and when a person should read the page first.
Crowdsourced transcription is the distributed production and review of text from digitised images by volunteers or a wider contributor community. The strongest documented pattern in heritage work is not volunteers or machines but a division of labour: handwritten text recognition (HTR) supplies a repeatable first pass across a homogeneous series, and people spend their scarce hours on review, adjudication, and the judgements a model cannot make — illegible passages, abbreviations, marginalia, named entities, editorial convention. The allocation decision turns on one question, asked per record series rather than per programme: is the machine draft good enough on this material that correcting it costs less than keying from a blank page?
That question has a different answer for a run of typed union correspondence than for a Tibetan manuscript collection, and a different answer again for nineteenth-century clerks' hands than for a fifteenth-century council register. Getting the routing right is most of the work.
What crowdsourced transcription actually is
The term covers several distinct task designs, and they are not interchangeable.
Microtask transcription
The page is broken into lines or fields and distributed, often with redundancy. Zooniverse's transcription task has volunteers collaboratively mark and transcribe lines, aggregating multiple independent readings.
Macrotask transcription
A contributor takes a whole page or document, usually with markup conventions, and the result is routed to staff or expert adjudication. Transcribe Bentham is the canonical case.
Consensus aggregation
Agreement is treated as the completion signal. The Library of Congress's By the People requires at least two volunteers to agree before a transcription is marked complete; Shakespeare's World used three or more independent transcriptions per word.
Blind double-keying
The stricter version: two people key the same source without seeing each other's answer, and disagreement flags a passage for adjudication. Note what it does and does not prove — it identifies where two readers diverged, not which reader was right.
Post-editing
Correcting a machine first pass is a different cognitive task from transcribing a blank page, and should be costed, trained, and quality-controlled as one. This is the point most hybrid pipelines get wrong.
Alongside these sits the metric vocabulary. Character error rate is the edit distance between system output and a reference transcription divided by the reference character count; word error rate applies the same calculation to words. OCR-D's quality-assurance specification defines normalised CER as a percentage from 0 to 100% and WER as the percentage of incorrectly recognised words. A score means nothing without its normalisation rules, segmentation, ground-truth method, sample, language, material and period — a point worth internalising before you read anyone's accuracy claim, including a vendor's.
Four things the headline programme figures don't tell you
Scale is not throughput
The Smithsonian Transcription Center reports 1,622,059 total pages and 108,341 "volunpeers" since June 2013. Trove covers over 20 million digitised newspaper pages, with more than 120 million OCR column-lines manually corrected by users as of March 2014 — described in the same documentation as only a small percentage of the collection. These are real achievements, and they are not a rate you can plan against. Participation is steeply unequal: in Transcribe Bentham's six-month test, seven "super transcribers" — 0.6 per cent of registered users — worked on 709 of 1,009 transcribed manuscripts. Plan around an active core of a dozen people, not a registration count.
Crowdsourcing is not free
Causer and Terras's evaluation of 4,364 checked and approved Bentham transcripts found staff checking averaged 207 seconds per transcript overall, falling to 141 seconds in the later period; 35 per cent of later-period transcripts needed no markup alteration, 47 per cent needed one to five changes, and 8 per cent needed ten or more. Crucially, the published cost projections exclude platform management, hosting, maintenance, upgrades, storage and long-term data management. Any comparison between volunteer and machine-assisted routes has to put those lines back in, alongside the rest of what a digitisation programme actually costs.
Accuracy is not a single number
By the People's dataset paper reports a 2019 sample of 240 characters from narrative Branch Rickey transcriptions at 98 per cent accuracy, and separately counts 703 character-level errors in 152,017 characters of Samuel J. Gibson material — a ratio of roughly 0.46 per cent, though the paper prints it as ".0046%", so the figure should not be quoted without doing the arithmetic. Neither is transferable to another hand, language, or set of conventions.
Consensus is not correctness
Redundancy suppresses idiosyncratic error. It does not catch shared error: everyone reading the same ambiguous surname the same wrong way, everyone silently modernising the same spelling, everyone dropping the same marginal insertion. That is why By the People pairs its consensus rule with staff spot-checking of at least 5 per cent of available datasets and full staff review of some collections.
What volunteers do best
Once you stop treating volunteers as free labour, it becomes clear what they are actually irreplaceable for.
Judgement about what cannot be read
Deciding that a passage is illegible, marking it as such, and resisting the urge to supply a plausible word is a skill. Machines are structurally bad at abstention.
Editorial convention
Whether to expand an abbreviation, how to record a deletion, what to do with a superscript insertion — these are project decisions that a person applies consistently and a model applies according to its training. Set them before a single page is distributed, in a documented transcription convention sheet.
Contextual and specialist knowledge
The Princeton case studies record Shakespeare's World volunteers identifying words absent from the OED and producing a ninety-year antedating of "partner" and a near two-hundred-year antedating of "white lie", alongside collective close reading on the project's discussion boards. No pipeline generates that.
Description and interpretation
Named entities, subject terms, and the connective knowledge that turns a transcript into a described record. Machine assistance exists here too, but the production-readiness varies sharply by sub-task.
Public engagement itself
For many institutions this is a mandate outcome, not overhead. A programme that turns a backlog into a community is delivering something a throughput figure cannot capture.
Where HTR should take the first pass
The case for machine seeding is strongest where the material is homogeneous, the volume exceeds any plausible volunteer supply, and the pages resemble what the model has seen. Large unindexed court and judicial series and climate data-rescue logbooks are the clearest examples: backlogs currently gated almost entirely on volunteer availability, where the substantive work is verification rather than first reading.
The Library of Congress's 2023 OCR button is the most candid published account of what seeding does. Volunteers choose whether to invoke it; the text appears in seconds and is editable like any other transcription; the team emphasises that OCR is imperfect and all text must be reviewed thoroughly. In the American Federation of Labor campaign — primarily typed letters — pages averaged 2.68 transcription actions overall against 1.38 for pages with OCR text. That is an action-count proxy in one campaign on largely typed material, and the team notes that volunteers sometimes invoke it on pages that are poor candidates, with mixed results.
The threshold effect is the thing to design around. In the Codex Runicus study, manual page transcription took 14–21 minutes while validating machine output took 5 or 7 minutes on the example given — but the authors note that correction time rises toward and past manual transcription as error density increases, and one configuration reporting 3.2 per cent CER still left over 20 per cent of characters missing. A draft with holes in it is not a draft; it is a proofreading trap.
Machine performance is also not uniform across periods and hands. The ICFHR 2016 READ competition, on German council minutes from 1470–1805 written by several hands, saw five systems report word error rates from 21 to 47 per cent. Models have improved substantially since; the lesson that survives is about variance, not vintage. Meanwhile, a peer-reviewed engine assessment found printed seventeenth-century Dutch and French Roman-type text reaching 0.82 per cent CER with language-model integration, while a Dutch States General set from 1576–1795 showed a 9.29 per cent base CER. Same century, same alphabet, order-of-magnitude difference.
One caution on reading those figures across: Roman-alphabet text in Dutch, French, German or English is not evidence about Latin-language material, and vice versa. Script and language are separate variables, and so is the hand.
Series to route away from the first pass
Send to human transcription, or to a controlled pilot, before committing a series: pages mixing dominant printed structure with dense handwriting, such as pre-printed ledger and deed-book forms, where recognition can favour the printed scaffolding over the manuscript entries; complex tabular layouts, where quality varies page to page; material in non-Latin scripts, where model coverage may not exist at all; and heavily damaged pages, where the first question is whether the problem is capture, recognition, or conservation.
The Tibetan pilot in Journal of Open Humanities Data is instructive on the alternative path: circa 100,000 pages, existing public models performing poorly, and a workflow built on iterative in-house model training with specialist annotator teams and triple-checking, reporting a Cohen's kappa of 0.66. That paper also notes plainly that training or fine-tuning your own models generally demands technical knowledge and server capacity — a real barrier for small institutions, and the reason many heritage teams default to volunteers by exclusion rather than by choice.
Choosing what generates the draft
If you seed, what you seed with determines whether volunteers are proofreading or being misled. This is the one comparison that matters at this stage, and it is worth stating flatly: a general chatbot is the wrong tool for a volunteer-facing first pass, because its errors are fluent. Garbled output announces itself and gets corrected. A confidently invented surname in period-plausible spelling does not, and it will pass consensus. The broader tool-class comparison is worth reading before you commit a series, but that single failure mode should settle the chatbot question for review workflows.
Leo is built for this stage. ATR-1 is a specialist transcription model that runs zero-shot — no per-collection model training, which removes the barrier the Tibetan team documents — and reads any language written in the Latin alphabet, handwritten or printed, including the early-modern founts, long s, ligatures and typographic abbreviation that conventional OCR normalises away. Non-Latin scripts are out of scope. It transcribes what is on the page: strikethroughs, insertions, marginalia, archaic orthography and editorial expansions survive into the draft rather than being smoothed, which is precisely what a reviewer needs in order to check text against image. Output showing failure patterns is withheld, retried, and refunded rather than served up as a plausible draft. Translation, where you need it, is a separate one-click Transformation writing to a new tab; the base transcription stays untouched, as do analytical outputs like summaries or named-entity extraction.
Around the model sits the workflow a programme needs: folders, per-document archival metadata, fuzzy search across a whole collection, TEI, Word, PDF and HTML export, and public read-only share links for publishing. The known weak spot is the one named above — pre-printed ledger and deed-book forms with dense handwriting. Corrections made in the app feed back into training, so the model improves across releases. For archives and academic projects, the Leo Transcription Grant offers up to 100,000 free credits in exchange for publishing the transcriptions and images openly within 24 months; applicants need to hold or clear the rights to publish, and printed material is not eligible — the free tier or a paid plan covers that.
Designing the hybrid, and proving it works
The operational pattern that holds up across the published cases is consistent: layout analysis and machine first pass; routing by legibility and material type; volunteer or specialist correction; tiered review with adjudication for disagreements; publication; then feedback into the model. Each stage constrains the next, in the same way the earlier stages of an archival digitisation workflow constrain everything downstream — and whether the end result is genuinely findable depends on decisions about full-text indexing made well before transcription begins.
Three things make it defensible. Build your own small ground-truth set from your own material before choosing an engine, and score against it — a published benchmark on someone else's corpus predicts very little about your registers. Reserve blind double-keying or targeted second review for high-stakes tokens: names, dates, numbers, sums, boundaries. And verify against the image, not against fluency, with the original displayed beside the text. If you are procuring the work rather than running it, the same discipline belongs in the contract as testable acceptance criteria with buyer-owned ground truth.
Several questions remain genuinely open. Whether machine seeding helps or hinders ordinary heritage volunteers has not been settled by a controlled independent test; the Library of Congress pilot is an institutional observation on largely typed material, and the rare-script study measured expert validation, not crowd behaviour. How often a plausible machine reading anchors a reviewer into confirming it — especially on abbreviations, names, and damaged passages — is not well quantified. Treat these as things to measure in your own programme rather than as settled findings.
What is not in doubt is that the allocation question deserves the care you would give to any other appraisal decision. Look at the series in front of you: how uniform the hand is, how much the layout varies, how damaged the paper is, how high the cost of a wrong reading, and how many people you can realistically count on next quarter. Then decide, series by series, what the machine drafts and what a person reads first. That judgement is the professional contribution — and it gets better the more of your own collection you have actually put through the process.
Frequently Asked Questions
What is crowdsourced transcription and how does it work?
Crowdsourced transcription is the distributed production and review of text from digitised images by volunteers or a wider contributor community. It takes several distinct forms: microtasking, where a page is split into lines or fields and distributed with redundancy; macrotasking, where one contributor handles a whole document under markup conventions before staff adjudication; consensus aggregation, where agreement between independent transcribers signals completion; blind double-keying, where two people key the same source and disagreements are flagged; and post-editing, where contributors correct a machine-generated first pass. These task designs are not interchangeable, and each carries different training, review and quality-control demands.
Is crowdsourced transcription actually free for archives?
No. Volunteer labour is unpaid, but the programme around it is not. Causer and Terras's evaluation of over four thousand approved Transcribe Bentham transcripts found staff checking averaged 207 seconds per transcript, falling to 141 seconds later in the project, with a minority of transcripts still needing ten or more markup changes. The published cost projections explicitly exclude platform management, hosting, maintenance, upgrades, storage and long-term data management. Any honest comparison between a volunteer route and a machine-assisted one has to put those lines back in before the arithmetic means anything.
When should HTR take the first pass instead of volunteers?
Machine seeding makes most sense where material is homogeneous, volume exceeds any plausible volunteer supply, and the pages resemble what the model has already seen well enough that correcting a draft costs less than keying from blank. Large unindexed court and judicial series and climate data-rescue logbooks fit that description: the substantive work is verification rather than first reading. Ask the question per record series, not per programme. Consider how uniform the hand is, how much layout varies, how damaged the paper is, and the cost of a wrong reading.
Does having multiple volunteers agree on a transcription guarantee accuracy?
No. Redundancy suppresses idiosyncratic error but does not catch shared error — everyone reading the same ambiguous surname the same wrong way, silently modernising the same spelling, or dropping the same marginal insertion. Blind double-keying has the same limit: it identifies where two readers diverged, not which reader was right. That is why the Library of Congress's By the People pairs its two-volunteer agreement rule with staff spot-checking of at least five per cent of available datasets and full staff review of some collections. Reserve targeted second review for high-stakes tokens: names, dates, numbers, sums, boundaries.
Which record series should be routed away from a machine first pass?
Four types warrant human transcription or a controlled pilot first. Pages mixing dominant printed structure with dense handwriting — pre-printed ledger and deed-book forms — where recognition can favour the printed scaffolding over the manuscript entries. Complex tabular layouts, where quality varies page to page. Material in non-Latin scripts, where model coverage may not exist. And heavily damaged pages, where the first question is whether the problem is capture, recognition or conservation. A low reported error rate is not sufficient reassurance: a draft with holes in it functions as a proofreading trap rather than a head start.