Rebuilding a 20-Year-Old Contract From a Phone Photo
A client sends a 20-year-old contract as a photo inside a PDF: fingerprint shadows, skewed pages, handwritten notes. Ordinary OCR gives up. This is how we turn it into a working document.
Every translation agency has seen this file. The contract is old, the original is long gone, and the only copy is a photo of the pages saved as a PDF. There are shadows from fingers holding the paper, the pages are tilted, and someone added notes by hand in the margins years ago. Standard OCR programs either fail completely or return text full of broken characters.
How we rebuild it
- Correct the geometry first: straighten skewed pages and fix the perspective of photographed sheets.
- Run professional OCR that keeps the document structure, including columns, tables and footnotes.
- Proofread by hand for merged or misread characters, which automatic recognition always leaves behind in poor scans.
- Rebuild the layout either one-to-one with the original or as a clean new design, whichever the client prefers.
Why it matters for translation
A CAT tool can only segment real text. If the source is a picture, linguists end up retyping it or working from a messy OCR dump, and errors in the source become errors in the translation. A properly rebuilt file gives the agency a document it can quote, translate and deliver like any other.
Not all scans are equal, but almost any of them can be recovered. The effort goes into the file before translation, so the linguists can spend their time on the language.
Key takeaways
- Photographed documents need geometry correction before OCR
- Good OCR keeps columns, tables and footnotes, not just words
- Poor scans always need a manual proofreading pass
- The layout can match the original one-to-one or be redesigned
- A rebuilt source file saves linguists from retyping