Anyone who has pulled a 19th-century parish record, a soldier’s field letter, or a great-grandmother’s diary out of an archive knows the feeling: the page is right there, but it might as well be in code. Old scripts don’t just have messy handwriting — they follow conventions that no longer exist. Letterforms have changed, words are abbreviated in ways nobody uses anymore, and even the spelling belongs to another era.
This is where general AI transcription gets genuinely hard, and where it pays to prompt deliberately. This guide focuses on historical hands — German Kurrent and Sütterlin, old English secretary and copperplate, and the archaic conventions that come with them. It’s a deep-dive companion to the complete guide to transcribing handwritten scans; the scan-quality and workflow basics there apply here too.
Know What You’re Looking At
A prompt works far better when it names the script, because that single fact tells the model which conventions to expect. A quick orientation:
Kurrent is the old German cursive used in everyday writing into the early-to-mid 20th century. Its letterforms look alien to modern readers — the e, n, and m in particular are easy to confuse, and a single word can look like a row of identical strokes.
Sütterlin is a specific, simplified form of Kurrent taught in German schools from around 1915 into the early 1940s. If your document is German and dates from that window, it’s very likely Sütterlin.
Secretary hand and copperplate cover much of older English-language material — secretary hand in earlier documents, the looping copperplate style in 18th- and 19th-century letters and registers.
You don’t need to be an expert. Even a rough guess (“German, probably Kurrent, around 1890”) gives the model a useful anchor, and you can refine it once you see the first results.
The Core Prompt for Historical Scripts
Name the script, name the era, and flag the specific traps the model should watch for. Replace everything in [brackets].
You are a paleographer specializing in [19th-century German Kurrent script].
Task: Transcribe the handwritten text in the image.
Context:
- Script: [Kurrent]
- Language and era: [German, ca. 1890]
- Watch for: the long s (ſ), easily confused e/n/m strokes, and
abbreviations marked with a superscript stroke or colon.
- Names likely to appear: [Wilhelmine, Bromberg, Westpreußen]
Rules:
- Reproduce the text in modern type, but keep period spelling exactly as
written (e.g. "thun", "seyn", "Thür"). Do not modernize.
- Transcribe the long s (ſ) as a normal "s".
- Expand an abbreviation only if you are certain. Otherwise transcribe it as
written and mark an uncertain expansion like this: u.[und?]
- Preserve original line breaks.
- Mark illegible passages as [illegible] and uncertain readings with [?].
- Do not guess at words you cannot read, and do not "modernize" the text
into contemporary phrasing.
The Watch for line does a lot of quiet work: telling the model which letters are confusable in this specific script primes it to slow down exactly where errors cluster. And as with any handwriting, the Names likely to appear list is your strongest single lever — historical place and person names are often spelled in forms that no longer exist on any map, so the model has no way to infer them.
Diplomatic vs. Normalized: Decide Before You Start
For historical material, there are two legitimate kinds of transcription, and they serve different purposes. Decide which one you want, because the prompts differ.
A diplomatic transcription reproduces the source as faithfully as possible — original spelling, abbreviations, capitalization, even line breaks. It’s what you want for scholarly citation, genealogical evidence, or anything where the exact original wording matters.
Produce a DIPLOMATIC transcription:
- Keep all original spelling, capitalization, and punctuation.
- Do not expand abbreviations; transcribe them as written.
- Preserve line breaks and the original line structure.
- Mark editorial uncertainty with [?] and gaps with [illegible].
A normalized (or reading) transcription modernizes spelling and expands abbreviations to produce something a contemporary reader can follow easily. It’s what you want for a family history you’ll share, or a readable edition.
Produce a NORMALIZED reading transcription:
- Modernize archaic spelling to current standard [German].
- Expand abbreviations silently (write the full word).
- Keep proper names in their original historical spelling.
- Note any expansion you are unsure about with [?].
A useful workflow is to generate the diplomatic version first as your ground truth, then ask for a normalized version from that transcription rather than from the image again — that keeps the two in sync and isolates the modernization step.
Calibrate With a Known-Good Example
Historical hands are usually consistent within one document — the same person, the same quirks, all the way through. That consistency is something you can exploit. If you can confidently read even a few lines yourself, transcribe them and hand the model the pair as a calibration example before the real task:
Here is a correctly transcribed sample from the same writer, to calibrate
to their hand:
[paste cropped image snippet]
Correct transcription: "[your verified transcription of that snippet]"
Now transcribe the full page in the next image using the same conventions.
Even one or two examples (few-shot) can lift accuracy noticeably when a single hand runs through the whole document, because the model learns this writer’s particular letterforms rather than guessing from the script in general.
When to Reach for a Specialized Tool
Be honest about the limits. For difficult or large historical corpora, purpose-built handwriting-recognition platforms — Transkribus is the best known in the archival and genealogy world — can outperform a general model, especially when you can train or apply a model tuned to a specific scribe or document collection. General AI models shine on one-off pages, mixed material, and cases where you want to lean on language understanding and context; specialized HTR shines on volume and on scripts where a trained model already exists.
The two aren’t mutually exclusive. A common pattern is to run a specialized tool for the bulk pass and bring a general model in for the passages it flags as uncertain, where contextual reasoning helps most.
Verify Names, Dates, and Numbers Against the Source
The same caution from all handwriting work applies with extra force here. Historical names, dates, and figures are exactly what genealogists and historians build conclusions on, and they’re exactly what the model is most likely to get subtly wrong — because it can’t infer them and may quietly “regularize” an unfamiliar spelling into a familiar one. Treat every name and date as something to confirm against the image, and keep the [?] markers visible in your working copy so you know which readings are still provisional.
A Note on Privacy
Much historical material is old enough that privacy concerns are minimal, but not all of it — 20th-century letters, medical records, and personnel files can still contain information about living people or their close relatives. If that’s the case for your documents, apply the same care you would with any sensitive scan: check that uploading is permissible, and prefer a service that doesn’t retain or train on your data.
Wrap-Up
Old scripts reward a prompt that names what it’s looking at. Tell the model the script and era, flag the confusable letterforms, supply the historical names, and decide up front whether you want a faithful diplomatic transcription or a readable normalized one. Calibrate with a verified sample when the hand is consistent, lean on a specialized tool when the volume justifies it, and confirm every name and date against the original.
For scan-quality tips, the base prompt, and variants for letters, forms, and tables, see the complete guide to transcribing handwritten scans with AI.

Leave a Reply