A PDF is a container, and it can hold either
The format has two ways to put words on a page. It can draw them as text, a list of characters with positions and a font, or it can place an image made by a scanner or a camera. Both look identical on screen and behave completely differently underneath.
Text can be selected, searched, reflowed and extracted. An image can only be looked at. When a scanner puts a page into a PDF, what lands in the file is one big picture per page, and a picture has no characters in it, no matter how clearly you can read them.
Three checks, none of them longer than a second
- Drag-select across a word. If nothing highlights, there’s no text under it.
- Press Ctrl+F and search for a word you can plainly see. A scan finds nothing.
- Zoom to 400%. Text stays crisp because it’s drawn at your zoom level; a scan goes soft and blocky because it’s being enlarged.
What happens when you tick OCR
Leave the option off and the extractor reads the text layer and nothing else: instant, exact, local. A scan comes back empty because there is nothing to read, and the result panel names those pages rather than handing you a document with silent holes in it.
Tick it and a recognition engine runs over the pages that have no text layer — only those pages. It looks at pixels and returns characters. The engine is about 7 MB and downloads the first time you use it, not when the page loads, so a visitor who never opens a scan pays nothing for it.
- Born-digital PDF
- Text comes out of the text layer. Layout is approximate, but the characters are exact.
- Scanned PDF
- Nothing without OCR. With it ticked, clean print comes out close; handwriting and faint faxes do not.
- Mixed document
- Typed pages come from the text layer. OCR touches only the image pages.
- Right-to-left or CJK text
- The text layer comes out, but the reading order of a complex line can be wrong. OCR is English-only, so a CJK scan stays empty.
What this OCR is good at, and what it isn’t
It reads printed English. A clean 300 dpi scan of a typed page, or a PDF exported from a picture of one, comes through well enough to search, copy and quote. Handwriting, faint faxes, skewed pages and anything in another language come back wrong or empty, and the result panel names the pages where it found nothing. Treat the output as a draft rather than a transcript and read it against the original, because a wrong digit in a reference number looks exactly like a right one.
It costs one download, once, and only if you ask for it: roughly 7 MB of recognition engine and English model, fetched the first time you tick the box. A visitor who never opens a scan pays nothing, and the document never leaves the tab. No server sees it, which is the entire reason this runs here instead of somewhere with a queue.
If the scan is bad enough that a human struggles
OCR isn’t the only step that gets harder as quality drops. A skewed page, heavy speckling from a photocopier, or a low-contrast fax will defeat most engines, and the ones that claim to handle it do so by guessing at characters.
What helps more than any tool is a better source. If there’s a paper original, scanning it again at 300 dpi in grayscale beats every post-processing option. If it came from someone else, asking them for the original digital file costs one email.
