Home/Compress/Scanned documents

Compress a scanned PDF

This is the case compression was built for. A scanned page is a large image, images can be re-encoded, and the difference between how a scanner saved it and how a JPEG can be saved is where the entire reduction comes from.

Drop the PDF you want to compress

or choose a file

Even dozens or hundreds of pages are fine. It's handled one page at a time, so the page never locks up

Why a scan has so much to give up

A scanner has no idea which parts of the page matter. It samples every square inch at high resolution in full color, and stores the result losslessly, so a page that’s 95% blank paper still costs the same per pixel as a page dense with text.

Re-encoding the page as a JPEG is lossy, and the thing it discards is mostly detail your eye can’t resolve at reading size. That’s why the reduction is dramatic here and negligible on a document the bank generated as text: on a text document there’s nothing redundant to remove, so re-encoding only adds.

What each resolution costs you in readability

The setting you choose is a straight trade between file size and how well small marks survive. This is what each one looks like on a scanned page.

Faithful: 200 dpi
Small print, stamps and signatures keep their edges. Use this if the file will be printed or inspected. The default for anything going to a lender or a lawyer.
Balanced: 150 dpi
Where most scans land. Reads fine on screen, prints acceptably, and often cuts the file by more than half again versus 200 dpi.
Strong: 110 dpi
Screen reading only. Numbers and small type start to blur at this point, so check a figure before you send one.
Extreme: 80 dpi
The largest reduction, and the one most likely to make a document unusable. Fine for a signature page, wrong for a table of figures.
Grayscale
Worth trying on a document scanned in color that has no meaningful color in it. A color-in-color-out scan of black ink on white paper carries two channels of pure noise.

The ceiling that quietly reduces your setting

Very large pages are scaled down on the way out to stay under the renderer’s pixel limit. In practice this means a 300 dpi setting on a big page may not come out as 300 dpi.

It isn’t a bug and not something you can switch off. The alternative is a canvas the browser refuses to allocate, which fails outright rather than degrading. It’s worth knowing because the figure in the result panel is what actually happened, and it’s the number to trust.

The result panel reports what happened, not what you asked for. Both routes are run and both sizes are shown, with the resolution that was actually applied. If the pixel ceiling trimmed your setting, the reported figure reflects the trimmed value.

A scan usually has no text layer to lose

The usual argument against re-encoding is that it destroys the text layer, form fields, links and bookmarks. On a scanned page most of that was never there. The page is a picture, so there’s no text layer to destroy.

Form fields are the exception worth checking. A scanned form with fillable fields drawn on top of it is a real and common thing, and re-encoding flattens them into the image. The compressor reports the field count it removed so you can see it happened.

Check for form fields before compressing a form. An AcroForm sitting on top of a scanned page is invisible on screen until you click it. If the recipient has to fill the file in, use the lossless route instead. The reduction is smaller and the fields survive.

What we can and can’t tell you about your file

There’s no table of real-world results here, and the reason is the same one that makes the site worth using: nothing is uploaded, so there’s no collection of customer scans to measure. A site showing averages from “thousands of real documents” either kept them or invented the number.

What you get instead is your own file’s two numbers, computed on your machine, in front of you. For a document this variable, where scan resolution, color depth and how much of the page is ink all change the result, that’s the only measurement that means anything.

Straight answers

How small will my scan get?

A full-color 300 dpi scan re-encoded at the balanced setting usually lands between a twentieth and a fortieth of its original size. A scan that was already grayscale and already compressed will give up much less, because there’s less waste in it to begin with.

Will the text become selectable if I compress it?

No, and nothing makes a scan searchable except OCR. If the file has no text layer now, compressing it won’t add one. It will remove less, because there was less to remove.

Can I tell it to skip the image route?

Yes, tick Keep the editable structure. On a scan that means the reduction drops to whatever the lossless route manages, which on an image-heavy file is usually very little. It exists for forms and for documents that will be machine-read.

The output is bigger than the input. Why?

Either the file wasn’t really a scan, or it was scanned grayscale at modest resolution and was already efficient. Both routes still run and the smaller result is the one you get, so a bigger output means both routes added weight, which happens on born-digital files.

Does compressing remove the metadata?

Yes, by default. On a scan the metadata is often the more revealing part: the scanner model, the operator name, sometimes the device serial. Worth leaving on unless the file is going nowhere.

More on compress

The full tool, with every option: Compress.