Repair a corrupt PDF — and find out whether yours can be repaired at all

Some damaged PDFs come back in a few seconds. Others are gone, and the most useful thing a tool can do is tell you which of the two you’re holding, instead of handing you a broken file and calling it fixed.

Drop the PDF you want rebuilt

or choose a file

The qpdf engine (about 430 KB) is fetched only when you press the button, and it runs in this tab.

Two kinds of damaged, and only one is fixable

The first kind still parses. The file loads in a reader, pages appear, and something underneath is slightly wrong: an object nothing points at, a cross reference table with a stale offset, a stream whose declared length doesn’t match what follows it. Readers tolerate these for years, and then one stricter program refuses.

The second kind is structurally dead. The file was truncated partway through a transfer, or the marker at the end that says where the table of contents lives is missing, or the object bytes themselves were damaged. A parser can’t get far enough to begin.

That distinction is the whole page. A rebuild is a second pass at writing a file that already parses. If the file can’t be parsed there’s nothing to write a second time, and no tool on your machine changes that.

What the rebuild actually rewrites

The tool parses the file from the beginning and writes it out again. The cross reference table is rebuilt, objects nothing refers to are dropped, and streams are recompressed. The content itself isn’t reinterpreted. Text stays text, and nothing is rasterised or re-flowed.

The side effect is usually a smaller file, because the debris of however many editing sessions went into it’s gone. On a 20 page document we measured 17.5 KB before and 12.8 KB after.

The four damages we tested against

Rather than describe the boundary in the abstract, we built four broken files and ran each one through this exact build.

  • A file truncated partway, so the final object never finished.
  • A file whose end marker points at an offset that no longer holds the cross reference table.
  • A file with damaged bytes inside an object.
  • A file whose stream declares a length that doesn’t match what follows.
All four were refused. The engine reads each one, decides it can’t proceed, and produces nothing. You get an explicit message rather than a damaged file with a success badge on it.

What a run tells you, in three states

The tool works in three steps and reports each one: a check of the file as it arrived, the rebuild itself, and a check of what came out. Every step ends in one of three states. No problems found, warnings only, or errors.

Warnings are worth more than a shrug. A file that rebuilds with warnings is a file that was repaired, but the warning names a structure the engine had to work around. If the same file keeps producing warnings, go back to the source and export it again.

Why you get a state rather than a sentence. The engine writes its own diagnosis to a console that a WebAssembly build can’t hand back to the page. Reporting the outcome accurately is possible. Reporting which object is wrong is not, so the page does the first and says so about the second.

What a repaired file is, and is not

File structure
Rewritten. Table rebuilt, unreferenced objects gone, streams recompressed.
Page content
Copied through, not re-encoded. Text remains selectable and searchable.
Fonts and layout
Untouched. Nothing is re-rendered, so nothing shifts.
File size
Usually smaller, sometimes identical. Growth is possible and isn’t an error.
Metadata
Kept as it was. Cleaning it’s a separate job.

When it says the file cannot be rebuilt

Then the damage is in the second category, and the honest answer is that no local engine recovers it. Not this one, not a paid one. Truncated files and files that have lost the pointer to their own structure are beyond repair by parsing, and parsing is what every repair tool does.

The way out is to stop repairing and regenerate instead. Ask whoever sent it to send it again, because an incomplete transfer is the single most common cause. If you hold the original document, export it once more. A file exported again from a working source isn’t a repaired file; it’s a correct one.

Be careful of the tools that promise otherwise. Truncated files aren’t recoverable by any of them, and the ones that take your file anyway tend to return a partial extraction presented as a repair. A repair that can’t be checked is worse than an honest refusal.

Straight answers

Is this the same as the paid PDF repair tools?

The mechanism is the same: parse the file and write it out again. If a paid tool succeeds where this one reports that it can’t, then that file was parseable and this one would have succeeded too. What you were paying for in that case is the diagnosis, not the repair.

Will rebuilding change how the document looks?

No. The pages are copied through rather than re-rendered. The text stays selectable, the fonts stay embedded, and the layout doesn’t move. What changes is the structure around the content.

It worked, and the file got smaller. Did I lose something?

You lost unreferenced objects and the leftovers of previous edits. Page content isn’t discarded to save space, so a large drop with no visible change is debris leaving the file. Compress the result if you want a further reduction.

Can it open a file that is password protected?

That is a separate job. If the file opens once you supply the password, its structure is intact and a rebuild will work on it. If it won’t open even with the password, the protection has to be dealt with first.

More on rebuild & diagnose

The full tool, with every option: Rebuild & Diagnose.