What DPI to Scan Old Documents: Resolution, Bit Depth, and the Specs That Decide Whether the Text Can Be Read

What DPI to scan old documents at, plus the bit depth, colour mode, compression, and measured image-quality settings that determine whether the text can be read.

Leo Team

September 17, 2026

Contents

This is a working guide to what DPI to scan old documents at, and to the capture settings that matter as much as the number itself — bit depth, colour mode, compression, and measured image quality. If you are digitising manuscript holdings once and expecting the files to serve readers, recognition models, and successors you will never meet, the specification you write now is the one you will live with.

For ordinary historical text — manuscript or print — capture the preservation master at 300 ppi or better, in colour or greyscale rather than bitonal, at 8-bit per channel minimum and 16-bit where the writing is faint. FADGI's Third Edition (May 2023) sets at least 294 ppi for Three Star on unbound manuscripts and rare materials and at least 396 ppi for Four Star, which additionally requires 16-bit; Metamorfoze specifies at least 300 ppi for handwritten material. The number alone guarantees nothing. What determines whether a page can be read — by a person or by a recognition model — is how many pixels land across a character, and whether the tone, sharpness, and colour of the capture are measurably good. Scan once, to the preservation target, and derive everything else from that file.

Why "what DPI" is the wrong question on its own

DPI is the term everyone uses and the term almost every scanner interface prints on screen, so it is worth being precise about what it names. Strictly, dots per inch describes a printer's or a scanner's output dots. PPI — pixels per inch — describes the spatial sampling of the digital image itself, which is what you actually care about. SPI (samples per inch) is the scanner-oriented equivalent. In preservation work the safest habit is to report captured pixel dimensions or measured sampling rather than the nominal figure the software claims.

That distinction matters because nominal ppi and optical resolving power are different things. Optical resolution is the detail the imaging system genuinely resolves. Interpolated resolution inserts calculated pixels between real ones — it grows the file and cannot restore detail that was never captured. The Library of Congress is explicit on this point in its still-image format preferences: pixel settings as of the time of creation are preferred over rescaled or interpolated settings. A 600 ppi file upsampled from a 200 ppi capture is a 200 ppi capture wearing a larger filename.

The second reason the question is incomplete: 300 ppi across a folio of large chancery cursive and 300 ppi across a tightly written eighteenth-century parish register produce very different amounts of evidence per letter. Resolution is a property of the page. Legibility is a property of the character.

Character height, not page resolution

The recognition-relevant measure is x-height in pixels — the height of the body of a lowercase letter such as x, measured in the captured image. The same ppi setting yields very different x-heights depending on type size, writing scale, and how densely the page is laid out.

For printed matter there is a rough conversion. A 12-point character is 12/72 of an inch tall, so its nominal pixel height is approximately ppi × 12/72 — about 50 pixels at 300 ppi, before accounting for the fact that the x-height is only a fraction of the point size. Handwriting has no fixed point-size equivalent at all; you have to look at the file and measure.

Vendor guidance offers a working number, though it should be labelled for what it is. A Nuance support page recommends 25–30 pixels for individual characters, noting this is met by a letter-size page of 12-point Times New Roman scanned at 300–400 dpi, and advises 400 or 600 dpi for type below 10.5 point. That is vendor material about modern print, not a controlled cultural-heritage study, and no controlled study establishes a universal x-height threshold for historical vernacular hands. Treat 25–30 pixels as a provisional floor to sanity-check against, not a standard to cite.

The practical procedure: open a representative capture at 100%, measure the x-height of the smallest hand in the series, and if it is thin — under roughly 20 pixels — raise the capture resolution for that series rather than for the whole project. Marginalia, interlinear insertions, and clerks' abbreviations are usually the smallest marks on the page, and they are usually the ones that matter.

What the standards actually specify

Current cultural-heritage practice is performance-based. FADGI's Third Edition aligns its star ratings with ISO 19264-style image analysis — Four Star nominally corresponding to ISO Level A, Three Star to Level B, Two Star to Level C, with One Star offered for reference use only and outside the standard. Metamorfoze works the same way, specifying sampling efficiency, MTF, tone, colour, exposure, gain, and noise tolerances rather than resolution alone.

Concretely, from FADGI 2023:

  • Unbound manuscripts and rare materials. Three Star: at least 294 ppi (300 minus 2%), 8- or 16-bit permitted. Four Star: at least 396 ppi (400 minus 1%), 16-bit required.
  • Oversize maps and posters. Three Star: at least 294 ppi. Four Star: at least 396 ppi. 8 or 16 bits permitted at both levels.
  • Microfilm. Three Star: at least 4000 ppi. Four Star: at least 4500 ppi, 8-bit greyscale.

The Three-to-Four-Star step is not simply more pixels. FADGI's reported tolerances for unbound manuscripts tighten across the board: tone response ΔL2000 below 3 becomes below 1.5; white-balance ΔE(ab) below 4 becomes below 2; lightness uniformity below 3% becomes below 1%; mean colour accuracy ΔE2000 below 3.5 becomes below 2; channel misregistration below 0.5 pixel becomes below 0.33 pixel. A rig that hits 400 ppi with uneven illumination and soft corners is not Four Star imaging.

Metamorfoze Version 1.0 (January 2012) sets at least 300 ppi for handwritten material, a minimum 85% sampling efficiency, and — as its worked example at 300 ppi — a minimum of 5 lp/mm at MTF10; claimed and obtained sampling rates may differ by no more than 2%. It permits 16- or 8-bit per channel at the full tier and 8-bit for Light and Extra Light, with noise limits of 16-bit STD ≤1024 and 8-bit STD ≤4.

One caution for anyone writing a specification: the Metamorfoze tiers do not map one-to-one onto FADGI stars. The relationship is a rough planning analogy, not an equivalence, and the widely circulated 2012 document has been superseded — verify against the current Dutch programme text before you commit it to an RFP or a vendor contract.

Bit depth: what it buys you, and what it doesn't

Bit depth is the number of tonal values encoded per channel. 8-bit greyscale carries 256 levels; 16-bit greyscale carries 65,536. In colour, 24-bit RGB means 8 bits per channel, 48-bit RGB means 16.

The Library of Congress preference is unambiguous: greater bit depth over lesser, 16-bit-per-channel data over 8-bit, 24-bit RGB over 8-bit indexed colour. The argument is capture latitude. Faint iron-gall ink over toned paper, a pencil diary entry, water-stained margins, show-through from the verso — these are all cases where the difference between text and background occupies a narrow slice of the tonal range, and more encoded levels means more room to separate them without banding or clipping.

What that argument does not establish is that every faded document needs 16-bit to be read successfully by a recognition model. The case for 16-bit is tonal latitude, not a demonstrated gain in transcription accuracy. So: choose 16-bit where FADGI Four Star requires it, and where you can see the writing sits close to the paper tone. Don't inflate an entire series to 16-bit on the assumption that the recognition step will reward you for it.

Colour, greyscale, and the bitonal trap

Colour records information about ink, paper, later annotations, seals, and degradation. Greyscale is appropriate where colour carries no evidential value. Bitonal reduces every pixel to black or white and discards weak or contextual information permanently — a faint correction, a pencilled marginal note, a passage of show-through that a human reader could disentangle but a threshold cannot.

The old argument for bitonal masters was storage cost, and that argument has largely expired. Bitonal is properly an access or derivative representation, produced from a richer master, and defensible as a capture mode only when the source is clean and the project specification explicitly permits it. If you are capturing manuscript hands, capture in colour or greyscale.

Compression follows the same logic. The Library of Congress prefers uncompressed bitmap data over lossy compression, listing uncompressed TIFF first among preferred colour and greyscale formats, and preferring lossless JPEG 2000 over lossy JPEG where JPEG 2000 is used. Lossy artefacts around letter edges are exactly the sort of noise a recognition model can misread as stroke evidence.

Master first, derivatives after

The organising principle behind all of this is that you capture once. IFLA's guidance on collection reproduction for preservation calls for a preservation workflow holding multiple properly packaged copies, including an archival master image and a production master image. NEDCC makes the preservation case directly: digital surrogates protect fragile and valuable originals from handling.

That second point is the operational one for anyone running a reading room. Every rescan is another exposure — another mount, another light cycle, another turn of a brittle leaf. Capturing to the preservation target the first time means the recognition step, whenever you run it and with whatever tool, works from a file that already exists.

It also means you should not tune capture to today's recognition engine. Tool guidance is uneven and tool-specific: Tesseract's documentation says it works best on images of at least 300 dpi; Amazon Textract's best practices ask for ideally at least 150 dpi, with a 15-pixel minimum text height noted as equivalent to 8-point text at 150 dpi; Google publishes pixel dimensions rather than a DPI rule; ABBYY offers automatic and manual resolution handling without publishing a numeric minimum. On the specialist side, Transkribus advises that around 300 dpi is fine and that larger images are not necessary, and Kraken asks for high-quality colour or greyscale scans, preferably at least 300 dpi, in a lossless format such as TIFF.

Read those together and the conclusion follows: there is no cross-engine threshold. Capture to the preservation standard, then derive whatever each tool wants — downsampled, converted, cropped — from the master. The derivative is disposable. The master is not.

If you are building this into a larger programme, resolution sits at the capture stage of a longer chain; the rest of that chain, from selection through QC to publication, is set out in our stage-by-stage archival digitization workflow, and the capture-side field technique for reading-room photography — lighting, geometry, focus, handling — is covered separately in the guide to photographing archival documents. For a broader view of how capture, description, and searchability constrain one another, see the archives, digitization and metadata workflows hub.

When resolution isn't the problem

Failed recognition runs are often blamed on DPI when the actual obstacle is elsewhere. It is worth separating the causes before you re-rig the copy stand.

Condition problems

Faded iron-gall ink, foxing, and show-through are not solved by more pixels. They are capture, conservation, or recognition problems in their own right, and the triage of faded and damaged documents is a separate diagnostic path.

Geometry and lighting

Skew, gutter curvature in tightly bound volumes, uneven illumination across the page, and channel misregistration all degrade recognition at any resolution. FADGI's tolerances exist precisely because these are measurable and controllable.

The script itself

Early-modern typography and idiosyncratic vernacular hands defeat general OCR even when nominal DPI is entirely adequate. The long s read as f, ligatures split or dropped, blackletter and Fraktur founts, typographic abbreviation with no representation in a modern character model, archaic orthography quietly "corrected" by a post-processing language model — these are model-fit failures, not sampling failures. That is a pattern practitioners report consistently rather than a finding from a controlled tool study, but it holds up in the reading room.

The third category is where a well-specified 400 ppi master and a disappointing text layer sit side by side, and it is the one worth naming clearly: no amount of additional resolution will make an engine trained on clean modern type read secretary hand.

Where the text layer comes in

Once the master exists, the question shifts from capture to recognition — and the two need to stay conceptually separate, because they fail for different reasons and are fixed by different means.

This is the stage Leo occupies. ATR-1 is a specialist handwritten text recognition model trained on images of historical documents, reading Latin-script material — English, French, German, Dutch, Spanish, Italian, Latin, and any other language written in that alphabet; non-Latin scripts such as Greek, Cyrillic, Hebrew, Arabic, and Indic and East Asian writing systems are out of scope. It reads printed matter as well as manuscript hands, including the early-modern founts and blackletter where conventional OCR struggles most, and it runs zero-shot: no per-collection model training, no segmentation step, no preprocessing pass before upload. It transcribes what is on the page — long s stays long s, abbreviation marks are not silently resolved, archaic spelling survives — rather than smoothing the text into modern prose.

On the resolution question specifically, two things follow from how it works. It weighs high-resolution visual evidence against textual context, which is the argument for feeding it the best derivative your master supports rather than a heavily compressed access copy. And because its errors are character- and word-level rather than fluent invention, they are checkable against the image displayed beside the transcription — which is the point of keeping a good master to check against. On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, at ATR-1's release, Leo scored roughly 5% character error rate — 61% fewer errors than the next-best model tested (Transkribus/Text Titan I ~13%, Claude Opus ~23.3%, Gemini 2.5 Pro ~24.8%, GPT-4.1 ~56.7%); the full comparison is published here. Pages that mix dominant printed structure with dense handwriting — pre-printed ledger and deed-book forms — remain the weaker case, and complex tabular layouts vary.

The broader question of what a recognition step can and cannot deliver from a given capture is worked through in the guide to making archives searchable, which separates finding-aid metadata from full-text transcription — two different deliverables with two different capture implications.

A working specification

For a typical manuscript series, in the absence of a project-specific mandate:

  1. Capture at 300 ppi minimum, 400 ppi where the hand is small, the page is dense, or the material is rare enough that you will not want to return to it.
  2. Measure, don't assume. Check sampling efficiency against claimed resolution, and check x-height in the captured file on the smallest writing in the series.
  3. Colour or greyscale, never bitonal at master. Colour where ink, paper, seals, or later annotation carry evidence.
  4. 8-bit per channel minimum; 16-bit where FADGI Four Star applies or the writing sits close to the paper tone.
  5. Uncompressed TIFF, or lossless JPEG 2000. No lossy compression in the master.
  6. Validate the whole image, not just the number — tone response, sharpness/MTF, colour accuracy, noise, illumination uniformity, registration.
  7. Derive access and recognition copies from the master. Never rescan for a tool.

Two of those steps are easy to skip and expensive to skip: measuring rather than trusting the nominal figure, and looking at an actual character at 100% before you commit a series to the copy stand.

The habit worth building is not memorising a number but knowing which measurement answers which question. Page ppi tells you what the rig delivered. X-height tells you whether the writing is legible. Tone and MTF tell you whether the capture is defensible. Bit depth tells you how much latitude you have left when the ink is faint. A specification that carries all four is one you can hand to a vendor, test on delivery, and stand behind in ten years — when the recognition tools have changed again and the master file is still the only thing you have.

Frequently Asked Questions

What DPI to scan old documents at?

Capture preservation masters at 300 ppi or better — 400 ppi where the hand is small, the page dense, or the item rare enough that you will not want to return to it. FADGI's Third Edition (May 2023) requires at least 294 ppi for Three Star on unbound manuscripts and rare materials, and at least 396 ppi with 16-bit for Four Star; Metamorfoze specifies at least 300 ppi for handwritten material. The number alone is not the whole specification: colour or greyscale rather than bitonal, 8-bit per channel minimum, and no lossy compression in the master matter just as much.

Is 300 or 600 DPI better for scanning old documents?

For most historical text, 300–400 ppi captured optically is sufficient; 600 ppi is worth it only when the writing is genuinely small or dense. What disqualifies a higher figure is interpolation — calculated pixels inserted between real ones, which grows the file without recovering detail that was never sampled. The Library of Congress prefers pixel settings as of the time of creation over rescaled or interpolated settings. A 600 ppi file upsampled from a 200 ppi capture remains a 200 ppi capture. Check optical resolving power and sampling efficiency rather than the nominal figure the scanner software reports.

Should I scan old documents in colour, greyscale, or black and white?

Colour or greyscale — never bitonal for a preservation master. Colour records ink, paper, later annotations, seals, and degradation; greyscale is appropriate where colour carries no evidential value. Bitonal reduces every pixel to black or white and permanently discards weak or contextual information: a faint correction, a pencilled marginal note, show-through a human reader could disentangle but a threshold cannot. The old justification for bitonal masters was storage cost, and that argument has largely expired. Bitonal belongs to access derivatives, produced from a richer master, not to capture.

What is x-height, and why does it matter more than DPI?

X-height is the height in pixels of the body of a lowercase letter, such as x, measured in the captured image — and it is the recognition-relevant measure, because the same ppi setting yields very different results across large chancery cursive and a tightly written parish register. Resolution is a property of the page; legibility is a property of the character. Open a representative capture at 100%, measure the smallest hand in the series, and if it falls under roughly 20 pixels, raise the resolution for that series. Marginalia and abbreviations are usually the smallest marks and often the most important.

Do I need 16-bit scans for faded documents?

Not always. The case for 16-bit is tonal latitude, not a demonstrated gain in transcription accuracy. 8-bit greyscale carries 256 tonal levels per channel; 16-bit carries 65,536, which gives more room to separate faint iron-gall ink, pencil, or show-through from a toned background without banding or clipping. The Library of Congress prefers greater bit depth over lesser. Use 16-bit where FADGI Four Star requires it, or where you can see the writing sits close to the paper tone — but avoid inflating an entire series on the assumption that recognition will reward it.

Share this article

© 2026 Leo Technologies Limited. All rights reserved