PDF OCR

A scanned contract, a photographed page, an old fax saved as PDF — all of them look like documents and behave like pictures: you cannot search them, cannot copy a line, and a screen reader has nothing to read aloud. This page recognises the words and puts them back where they are on the page, invisibly. The scan looks the same; everything else starts working.

Recognition is a guess, not a reading: a clean printed page comes out almost perfect, a phone photo of a crumpled receipt does not. Nothing is auto-corrected — a plausible fix to a number is a mistake you would never catch.

The scan is not uploaded: the engine and the language model are downloaded into your browser and work there. What comes out is your own scan with an invisible text layer on top — the picture looks the same, but search, copying and screen readers work.

Nearby: Image to Text · PDF to Text · PDF to Word · Compress PDF

How to use

  1. Choose the scanned PDF.
  2. Pick the language of the document — recognition depends on it more than on anything else.
  3. Press the button, wait, and download the searchable PDF or the plain text.

Good to know

What a searchable PDF actually is

It is your scan with a second layer: the recognised words, placed exactly over the printed ones and drawn with invisible ink. The page looks identical — same paper, same stamps, same handwriting in the margin — but Ctrl+F finds a name in it, a phrase can be copied, and a screen reader can read it aloud. That is why the picture is kept rather than replaced by clean text: a scan is often a document of record, and re-typesetting it would make it something else.

Recognition is a guess, not a reading

A clean printed page comes out nearly perfect. A photo of a crumpled receipt does not, and no engine changes that. So the average confidence is shown after the run, and nothing is auto-corrected: a plausible correction to a digit is the kind of error nobody ever catches. Read the text before you rely on it — especially the numbers.

Frequently asked questions

Is my scan uploaded?

No. The engine and the language model are downloaded into your browser and run there; the file stays on the device.

How much is downloaded?

Around two to three megabytes: the engine plus one language model. It is stated on the page before you press, and after the first run everything works offline.

My PDF already has text — should I run this?

No, and the page says so when it notices: recognition would replace good text with a guess. Use «PDF to text» instead.

How many pages can it handle?

Up to fifty. Beyond that a browser would spend tens of minutes, and the page says how many were done.

Related tools