Filozy

Guide · updated 1 September 2026

OCR a scanned contract on your own computer

OCR adds an invisible text layer to a scanned page so it can be searched and copied. On Filozy the recognition engine runs in your browser, so the scan never leaves your device. Accuracy depends on resolution, contrast and skew; verify names, dates and totals against the page before relying on them.

Tools used:OCR PDFAuto crop PDFRedact PDF

A scanned contract is a photograph of a contract. You can read it, but your computer cannot: no search, no copy-and-paste, no way to find the indemnity clause across forty pages except by eye. OCR — optical character recognition — fixes that by reading the picture and adding the words as an invisible layer over the page.

The catch is that the documents people scan are exactly the ones they should not upload. Contracts, IDs, medical records, signed statements. Most OCR happens on someone else’s server. This guide does it on yours.

What OCR produces

Running OCR on a scanned PDF gives you two things:

Nothing is “converted”. The scan remains the scan; the text is an addition.

What decides accuracy

OCR is statistics, and its confidence rises and falls with the input:

You cannot fix a bad scan with a better algorithm. Rescanning at 300 DPI, flat, in good light, is worth more than any setting.

Step by step

  1. Trim and straighten first. Open the scan in Auto crop to remove scanner borders and uneven margins. Clean edges help the engine find the text block.
  2. Open OCR PDF and choose the trimmed file. The recognition engine is downloaded from Filozy’s own servers at this point — it is program code, not your document going the other way.
  3. Run the tool. Progress and a confidence figure are shown as it works; a low figure on a page is a hint to look closely at that page later.
  4. Save the searchable PDF, and the text file if you want it.
  5. Verify before you rely on it. Search the result for a clause you know is there. Compare every number, date, party name and defined term you will act on against the page image. OCR errors cluster in exactly those places: 0 and O, 1 and l, decimal points, currency symbols.

What to do next

The searchable PDF is the input to everything else:

A ready-made sequence for this — trim, recognise, number, protect — exists as the Prepare a scanned contract for filing workflow, which carries the file from one tool to the next without re-selecting it.

Honest limits

Questions people ask

Does OCR change how the page looks?

No. The scanned image stays exactly as it was. OCR adds an invisible layer of text positioned over the words in the picture, which is what lets you search and select them. Print it and it looks identical.

Which languages does Filozy's OCR recognise?

English. The engine's language data is served from Filozy's own servers when you run the tool, and only English is included today. Documents in other languages will be recognised poorly.

How long does it take?

Roughly a second per page on a laptop and several seconds per page on a phone, after a one-time download of the engine for the session. The first page is the slowest. Keep the tab in the foreground on mobile so the browser does not pause the work.

Related guides