Looking for DT Photo?
Select Page

Handwritten, faded, and long-forgotten text doesn’t stay locked up in an image file. It becomes searchable, translated, readable text.

A finished capture is a starting point, not an end point. A high-resolution photograph of a handwritten letter is, to a keyword search, indistinguishable from a photograph of a blank page — the text is there, but it isn’t data yet. DT Cipher is DT DigiLabs’ advanced OCR and translation product, built specifically for the kind of text that breaks ordinary OCR: cursive hands, period typefaces, faded or damaged originals, and languages consumer transcription tools were never trained on.

What Makes It Different

General-purpose OCR is trained on modern printed text — clean fonts, standard layouts, one language at a time. Historic material breaks almost every one of those assumptions at once: handwriting styles that have fallen out of use, ink that has faded unevenly, paper that has degraded, and vocabulary and abbreviations a contemporary language model has never encountered. DT Cipher is built around three things that matter specifically for this kind of material:

  • Collection-specific vocabularies. Instead of relying on general-purpose language patterns, DT Cipher builds controlled, topic-specific vocabularies around what a given collection is actually about — the names, places, terminology, and abbreviations that appear again and again in that particular archive — which measurably improves accuracy over a one-size-fits-all model.
  • Lost and endangered-language capability. Beyond mainstream multi-language output, DT Cipher is applied to the harder end of the problem: historic, endangered, and under-represented languages like Ottoman Turkish, Ancient Greek, or Norse, where getting the transcription right the first time matters because there may be nobody left to catch a subtle error.
  • Delivery in the formats institutions already use. Output isn’t a wall of plain text — DT Cipher delivers searchable text alongside your images in PDF, PDF/A, METS/ALTO sidecar XML, and plain text, so it drops into existing digital collection systems instead of requiring a new one.

What It Does

  • Next-generation OCR built for historic material — cursive handwriting, faded ink, esoteric period typefaces, and damaged or degraded documents
  • Multi-language output, including translation and recovery of lost or endangered languages
  • Controlled, topic-specific vocabularies built around your collection’s actual subject matter to improve accuracy
  • Delivers searchable text alongside your images, in the formats your systems already use — PDF, PDF/A, METS/ALTO sidecar XML, and plain text

Who It’s For

  • Libraries and archives with historic correspondence and manuscripts
  • Government and legal institutions with historic records
  • Religious and cultural institutions with sacred or liturgical texts in historic scripts
DT Cipher providing transcription and translation of historical handwritten text

Typical Use Cases

Recovering a difficult historic script

DT DigiLabs is currently applying DT Cipher to a collection of vernacular Yiddish letters handwritten in cursive Ashkenazi Hebrew script — work that requires more than OCR alone. It requires a genuine, collection-aware understanding of a script that is easy to misread even for a trained human reader, let alone a generic transcription tool. That’s the kind of project DT Cipher is built for: not high-volume printed text, but the hardest individual documents in a collection, where accuracy depends on real linguistic and paleographic expertise built into the model rather than a one-size-fits-all OCR pass.

Making a correspondence collection searchable

Institutional and personal correspondence collections are often the most requested and least searchable material in an archive — hundreds or thousands of handwritten letters that currently require a researcher to open and read each one to know what’s inside. DT Cipher turns that into full-text search: a researcher can surface letters mentioning a name, place, or event without opening a single folder first.

Translating records nobody on staff can read

Government, legal, and religious archives frequently hold material in languages or scripts no one currently on staff can read fluently — colonial-era records, historic liturgical texts, or documents in a regional language that has since fallen out of everyday use. DT Cipher provides a first-pass transcription and translation that gives staff and researchers a working entry point into material that would otherwise require an outside specialist for even a basic summary. Such initial translation can help identify which parts of the collection most urgently warrant the time of a professional human translator.

RESPONSIBLE AI: DT Cipher’s transcriptions and translations are AI-assisted first drafts, built to be reviewed by a qualified linguist or archivist before publication — especially for lost or endangered languages, where accuracy is never optional.

Works on What You Already Have

DT Cipher can be run on the high-fidelity raw captures produced during a DT DigiLabs digitization project — but that isn’t required. It works on high-resolution images your institution already has — from any digitization effort, old or new — which makes it a practical way to unlock the value in a manuscript or correspondence collection that’s already been scanned but never transcribed.

Pairs Well With

  • Digitization & Capture — start with a fresh capture, then transcribe and translate
  • DT Registrar — combine searchable text with object and subject metadata for full collection search

If your institution has historic text nobody has had the time — or the language expertise — to read closely, DT Cipher can help. Talk to our team about your collection.

Related Resources