Home/Strip Metadata/Hidden fields

What your PDF remembers about you

The page is the part you wrote. Underneath is a second document, the one the software kept about itself, and it’s usually more specific about you than anything on the page.

Drop in a PDF and see what it's carrying

or choose a file

Read-only analysis; your original file is never modified

Two places the details hide

The Info dictionary is the obvious one: title, author, subject, keywords, the producing application, and two timestamps for created and last modified. It’s a short list of plain text fields, and every reader can show it to you.

XMP is the one people don’t know about. It’s an XML block that can sit alongside the pages and record far more: the full editing history, the original file path, the camera or scanner model, sometimes coordinates. It’s often larger and more revealing than the Info dictionary, and nothing in a normal reader surfaces it.

The dangerous field is rarely the one you would guess. A title of “Invoice” tells nobody anything. An author field reading your full name, a producer naming the exact software version your firm runs, and an XMP block carrying the path to your desktop is a different matter, and it travels with every copy of the file.

What each field actually gives away

Author
The account name the software was licensed to. On a work machine this is often your full name, not the name on the document.
Creator
The application that authored the content: Word, InDesign, a scanner driver. Reveals the workflow behind the file.
Producer
The library or virtual printer that produced the PDF, usually with a version number. Combined with Creator, it identifies the toolchain.
Creation and modification dates
A timeline. When the file was made and when it was last touched, which isn’t always when the content stopped changing.
Custom and company fields
Whatever the originating system chose to write. Templates often stamp a firm name, a client code or a matter number here without anyone deciding to.
XMP history
The most revealing block when present. Original path, revision history, sometimes device information.

Where the risk actually sits

Metadata is a small leak, and it matters most in the cases where the document is going somewhere adversarial. A CV sent to a company you don’t want knowing your current employer. A redacted document where the author field names the person the redaction was hiding. A quote going to a competitor, carrying the version of the software that produced it.

It also matters cumulatively. One file with an author field isn’thing. A hundred files from the same source, each carrying the same fields, is a pattern, and it’s the pattern that gets noticed, not any single file.

See it before deciding

The tool reads the file and shows you the whole picture: the Info fields it found, the risk rating it assigns each one, and whether an XMP block is present. Nothing is modified at this stage, because it’s read-only and the original file is untouched on your disk.

Then choose one of two operations. Wipe everything removes the Info dictionary entirely and deletes the XMP block from the document catalog. Write specific fields lets you set the ones you want and clear the rest.

In rewrite mode, an empty field is cleared, not preserved. That’s deliberate. If leaving a field blank meant “keep the old value”, the old value would survive, which defeats the purpose of opening the file to change it.

How you know it worked

Most tools tell you the metadata is gone. This one reopens the document it just wrote and walks its structure to confirm two things: the Info dictionary reference no longer exists in the trailer, and there’s no XMP stream in the catalog.

That’s a structural check rather than a text search, which matters more than it sounds. Metadata inside a compressed object stream is invisible to a plain text scan, so a tool that proves its work by grepping the output can report clean on a file that isn’t. Confirming the reference is gone doesn’t depend on how the bytes were stored.

Straight answers

Does the file I am looking at get changed?

Not by the analysis. The tool reads and reports. An export writes a new file, and the one you opened is left exactly as it was.

Why did my file arrive with the previous author’s name in it?

Because the template was created on someone else’s machine and the field was never overwritten. It’s one of the most common forms of this problem: a document that has been through five hands and still carries the first one’s name.

Does clearing metadata change the pages?

No. The pages are copied through as they were. Only the document-level fields and the catalog’s XMP reference are removed.

Can it read an encrypted PDF?

It can read the Info dictionary of a file with an owner password, since the content isn’t encrypted. A file with a user password can’t be opened at all without the password, which is a format limit and not a choice made here.

Is there anything metadata doesn’t cover?

Yes, and it’s worth knowing. Text you can read on the page isn’t metadata. If a name or an account number is printed on the page, stripping metadata won’t touch it, and that’s the redact tool’s job.

More on strip metadata

The full tool, with every option: Strip Metadata.