Adding Transcription to Tropy: Reading the Photos You've Already Organized

Explains that Tropy manages and describes research photos but does not transcribe them, and walks through adding an HTR step (comparing OCR, general LLMs, specialist HTR, and Leo) to a Tropy-organized collection.

Leo Team

July 26, 2026

Adding Transcription to Tropy: Reading the Photos You've Already Organized

This is a practical guide to Tropy transcription — what Tropy does, what it does not, and how to add a reading step to a collection you have already organized. If your research images sit sorted and described in Tropy but the text inside them is still locked in the pictures, this explains where the boundary falls and how to cross it without losing the organization you built.

If you have organized your archive in Tropy and now want the text of those photographs — searchable, editable, citable — you need a transcription tool alongside Tropy, not inside it. Tropy is a photo-management application from the Roy Rosenzweig Center for History and New Media. It stores, describes, and orders your research images well, but it does not read them. Its transcription pane holds text you type by hand. To turn a folder of manuscript images into full text, you export or upload those images to a handwritten text recognition (HTR) engine, transcribe there, and bring the results back into your workflow. What follows works through exactly where that boundary falls, and how to add the reading step without abandoning the organization you already have.

Why Tropy organizes but does not transcribe

This is the most common misconception about Tropy, and it is worth dispelling precisely, because the confusion costs people weeks. Tropy has a pane historically labelled for transcription, and the label leads researchers to assume the software is doing optical work on the image. It is not. According to Tropy's own documentation, you transcribe or take notes by clicking in the bottom pane of the item view "and begin typing." The text autosaves as you type. That is the whole mechanism: a place to keep the words you have already read, not an engine that reads them for you.

The consequence matters for search. Inside Tropy, what is searchable is metadata and the notes you have manually entered. The photographs themselves stay images. A folder of two hundred deed-book spreads is, to Tropy's search, two hundred opaque pictures with whatever descriptive fields you have filled in. The moment you want to find every occurrence of a surname across the run, or locate the one page that mentions a particular parcel, you have hit the organize-versus-read gap. Tropy has shortened the path from finding sources to writing about them, exactly as it advertises. It has not read them.

Tropy's design is deliberate and, on its own terms, correct. It is built around customizable Dublin-Core-based metadata templates and a plugin architecture used chiefly for export. Over the years the community has raised native OCR and HTR on the forums repeatedly, and the plugin conversation has explored letting Tropy back "transcription" objects with an external engine. But whether any maintained plugin now produces machine transcription — as opposed to bridging metadata — is not clearly established. In practice, the reliable pattern is to treat transcription as a separate stage and Tropy as the organizational hub it was designed to be. This is the honest shape of a manuscript research pipeline: Tropy for capture and description, a dedicated engine for reading, and export standards to move text between them. It is also the larger question of which tool fits where in a manuscript workflow — a question worth settling before you commit a collection to any one path.

What "adding transcription" actually requires

Adding transcription to a Tropy-based workflow is really three sub-decisions, and it helps to name them before reaching for a tool.

How the images leave Tropy

Tropy exports metadata and photos — to JSON-LD, to a zip archive, to Omeka via plugin. For transcription you mostly need the images themselves in a common format: JPG and the other standard image types Tropy already holds. You do not need Tropy to hand off anything clever. You need your image files, which you already have, in a folder you can point a transcription engine at.

What engine reads them

This is the decision that determines whether you get usable text or a plausible-looking mess. It is not a small choice, and the rest of this article is mostly about it.

How the text comes back

Once transcribed, where does the text live? Some researchers reimport into Tropy's notes field. Others keep transcription and analysis in the transcription environment and treat Tropy purely as the image library. Round-trip interoperability — whether TEI or PAGE XML from a specialist engine imports cleanly back into Tropy's notes without manual conversion — is not clearly documented, so plan for some copy-and-paste or a settled division of labor rather than a seamless loop. The pragmatic answer for most individual researchers is to let text live wherever they will actually search and edit it, rather than insisting on a single round-trip.

Of these three, the middle one is where research stands or falls. So the substance of "adding transcription" is choosing the reader.

Choosing the reader: what breaks, and what to use instead

The instinct, once you have your images out of Tropy, is to reach for whatever is nearest — a general OCR service, or lately a chatbot. For historical manuscripts this is where the trouble begins, and it is worth understanding the specific failure so you can avoid it rather than merely be warned off it.

Why general OCR is the wrong reader for old hands

General OCR — Google Cloud Vision, Amazon Textract, ABBYY FineReader — is engineered around clean, modern, high-contrast print with regular letterforms and predictable spacing. On a modern page it is fast, cheap, and accurate. The distinction that matters is OCR versus HTR: OCR targets machine-printed text, HTR targets handwriting, and the assumptions baked into an OCR pipeline are precisely the ones historical material violates. A Berkeley Library comparison of OCR tools on nineteenth-century documents found ABBYY strong on printed paragraphs but struggling with anything off the expected grid. On manuscript hands, general OCR does not degrade gracefully; it produces noise.

Even historical print trips it — the long s read as f, ligatures split or dropped, blackletter and Fraktur type it has never modeled, uneven inking and show-through read as character evidence. So if your Tropy collection is early-modern books or pamphlets rather than handwriting, general OCR is still the wrong instrument, for the same underlying reason.

Why a general LLM is the dangerous reader

The newer instinct is to paste a page into ChatGPT, Claude, or Gemini. These models will return fluent, confident text, and that fluency is exactly the hazard. ChatGPT struggles on historical handwriting for structural reasons: the models downsample the high-resolution image the reading depends on, and they were never trained for this ultra-low-resource task. What they produce when uncertain is not a garbled character you can catch against the image — it is a plausible fabrication. The evidence here is genuinely mixed and should be read with care: a University of Virginia Library evaluation found Gemini the most accurate among general AI tools tested on primary sources, and one practitioner roundup argues such tools now approach usable accuracy on eighteenth- and nineteenth-century hands. But these are single-corpus, non-peer-reviewed assessments, still requiring line-by-line verification, and other sources document exactly the fabrication risk that makes LLM output hard to check. Fluent, plausible errors are the most dangerous failure mode precisely because they read as correct. A misread character advertises itself; an invented clause does not.

Specialist HTR: the reader built for the material

The tools actually built for this are specialist HTR engines — Transkribus, eScriptorium on the kraken engine, OCR4all for historical printings. These read old hands where general tools cannot. The historical friction has been the training step: the field's public-model hubs cover English, French, German, Italian, and Dutch across several centuries, which lowers the entry cost considerably, but for a niche hand you may still be curating ground truth and training a model before you get clean output. Accuracy, when a well-matched model exists, is strong: Transkribus describes a good CER for handwritten text as 2–8%, and a 2024 Journal of Documentation study reported CERs as low as around 1% on a French-language model. If you are weighing these against each other in detail, the broader comparison of handwriting transcription software is the place to do it.

Where Leo fits into a Tropy workflow

If you have organized in Tropy and want transcription that reads Latin-script manuscript hands out of the box — no model-training detour — Leo is built for exactly this stage. It sits where Tropy stops. You take the images you have already sorted and described, upload them, and transcribe. ATR-1, Leo's specialist model, is zero-shot — ready on upload, with no page-after-page training step to complete before you see usable text. That single difference is what makes it a comfortable addition to a Tropy pipeline rather than a project in itself.

On accuracy where it can be measured against the alternatives head to head: on a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library at ATR-1's release, Leo scored roughly 5% character error rate — 61% fewer errors than the next-best model, with Transkribus's Text Titan I at about 13%, Claude Opus about 23.3%, Gemini 2.5 Pro about 24.8%, and GPT-4.1 about 56.7% (full benchmark data). That is one corpus and one language; treat it as evidence for early-modern English hands, not a universal claim. Leo reads any language written in the Latin alphabet — English wills, French notarial records, Dutch registers, German parish books — with performance strongest in English and strong across the other major European languages. It is the writing system that defines the scope, not the language: non-Latin scripts (Greek, Cyrillic, Hebrew, Arabic, and East Asian systems) are out of scope. And transcription is not translation — Leo transcribes what is on the page; rendering a foreign-language page into English is a separate one-click Transformation that writes to a new tab, leaving the base transcription untouched.

Two things make it fit the Tropy user specifically. First, source integrity: ATR-1 is trained to transcribe what is on the page rather than to smooth it into modern prose — strikethroughs, marginal additions, archaic spelling, and editorial expansions survive rather than being silently normalized. It is engineered against the fluent-fabrication failure mode too: output showing failure patterns is hidden, retried, and the credit refunded if it cannot succeed, so the errors that reach you are the recoverable, character-level kind you check against the image beside the text. Second, the organizational overlap. Leo brings a Tropy-like structure of its own — custom folders, per-document Dublin-Core-adjacent metadata, the original image shown side by side with the transcription — plus global fuzzy search across every transcription, which is exactly the full-text search Tropy cannot give you. Export runs to Word, PDF, HTML, and TEI XML.

Whether you keep Tropy as your image library and treat Leo as the reading-and-search layer, or migrate the collection wholesale, is a judgment call about your own habits; the side-by-side of where each tool fits in a manuscript workflow works through that trade-off in more detail. If your material is historical print rather than handwriting, Leo reads that too — but note that printed matter is not eligible for Leo's Transcription Grant, so the free tier or a paid plan is the route there.

The step Tropy always leaves for you

Whichever reader you choose, one part of the work does not move. A machine transcription is a first pass, not a finished source. The reason to keep the original image beside the text — the reason Tropy shows it, and Leo shows it — is that verification is where scholarship actually happens: checking the machine against the image, giving names, dates, and figures the closest attention, and applying the editorial judgment no model supplies. Automation changes what that judgment costs you, not whether you owe it. Tropy organized the photographs; a transcription engine read them; the reading that turns text into evidence is still yours.

Frequently Asked Questions

Does Tropy transcribe images automatically?

No. Tropy is a photo-management application that stores, describes, and orders research images, but it does not read them. Its transcription pane holds text you type by hand — you click in the bottom pane of the item view and begin typing, and the text autosaves as you go. That is the whole mechanism: a place to keep words you have already read, not an engine that reads them for you. To turn manuscript images into full text, you export the images to a dedicated handwritten text recognition engine, transcribe there, and bring the results back into your workflow.

Why is only some of my Tropy collection searchable?

Inside Tropy, what is searchable is metadata and the notes you have manually entered — the photographs themselves stay images. A folder of two hundred deed-book spreads is, to Tropy's search, two hundred opaque pictures with whatever descriptive fields you have filled in. The moment you want to find every occurrence of a surname across a run, or locate the one page mentioning a particular parcel, you hit the organize-versus-read gap. Full-text search requires transcribed text, which means running your images through a transcription engine first, then searching that text wherever it lives.

Can ChatGPT transcribe old handwriting reliably?

No, not reliably. General LLMs like ChatGPT, Claude, and Gemini return fluent, confident text, and that fluency is the hazard. They downsample the high-resolution image the reading depends on and were never trained for this ultra-low-resource task. When uncertain, they produce plausible fabrications rather than garbled characters you can catch against the image. A misread character advertises itself; an invented clause does not. Evaluations are mixed and single-corpus, still requiring line-by-line verification. Fluent, plausible errors are the most dangerous failure mode precisely because they read as correct.

Why does general OCR fail on historical documents?

General OCR — Google Cloud Vision, Amazon Textract, ABBYY FineReader — is engineered around clean, modern, high-contrast print with regular letterforms and predictable spacing. The assumptions baked into an OCR pipeline are precisely the ones historical material violates. On manuscript hands it does not degrade gracefully; it produces noise. Even historical print trips it: the long s read as f, ligatures split or dropped, blackletter and Fraktur type it has never modeled, uneven inking and show-through read as character evidence. OCR targets machine-printed text; handwriting needs HTR, a different technology built for the material.

How do I add transcription to a Tropy workflow with Leo?

Take the images you have already sorted and described in Tropy, upload them to Leo, and transcribe. Leo's specialist model, ATR-1, is zero-shot — ready on upload, with no page-after-page training step before you see usable text. It reads any Latin-alphabet language, strongest in English, and preserves what is on the page: strikethroughs, marginal additions, and archaic spelling survive rather than being normalized. Leo brings its own custom folders, per-document metadata, side-by-side image and text, and global fuzzy search across every transcription. Export runs to Word, PDF, HTML, and TEI XML. Keep Tropy as your image library or migrate wholesale — that is your call.

© 2026 Leo Technologies Limited. All rights reserved