Home/Extract Text/PDF to Markdown

Turning a PDF into Markdown

A PDF knows where marks are, not what they mean. Markdown is entirely about meaning. Conversion is therefore a set of guesses, and it’s worth knowing which ones before you start editing the output.

Drop the PDF you want the text from

or choose a file

Reports, papers, e-books, invoices. Scanned pages have no text layer; tick "OCR scanned pages" below to read them.

What comes across, and what doesn’t

The whole list, so nothing in the output comes as a surprise.

  • Comes across: the words, their order, and the page they were on.
  • Comes across: headings, if you turn the guess on. It compares each line against the median body size and writes one hash for a line clearly larger, two hashes for a moderate step up. Two levels, no more.
  • Doesn’t come across: bold, italic and inline emphasis. The output has no way to know which words were emphasized, and inventing emphasis from font weight produces a file that looks structured and isn’t.
  • Doesn’t come across: lists, tables, links, footnotes, and captions as captions. A table arrives as its rows in reading order, which is often enough to rebuild by hand and never enough to do automatically.
  • Doesn’t come across: images and drawings. Markdown can reference a file, but a PDF conversion has no image to point at.

How the heading guess actually decides

Every line carries a font size. The tool takes the median across the page as the body size, then compares each line against it: a large step up becomes a level-one heading, and a smaller step becomes level two.

It’s a guess by construction, so it will promote an oversized caption now and then. The result panel reports how many lines it promoted, which is the number to check against the document in front of you.

Turn the guess off for anything with a designed first page. Cover pages, adverts and magazine layouts use size for effect rather than for hierarchy, so the guess will produce headings that aren’t headings. The output is still editable, so clear the hashes you don’t want.

Page boundaries are kept on purpose

The Markdown export writes the file name as a title, then one section per page, then that page’s lines. A heading per page sounds like noise until you need to cite something: a page number you can see in the file is worth the extra line.

Turn it off for documents where pages are an artifact of printing. A web page someone exported to PDF has no page structure worth respecting.

Repeated headers and footers

A running header on every page becomes a paragraph on every page. The filter drops a line when it repeats on at least 60% of pages, judging only the first three and last three lines of each page, the bands where running heads live.

Because it’s a threshold rather than a rule, it will occasionally keep a header, or drop a line that happened to repeat. The result panel says what it removed, so the check is a glance rather than a reading.

The output is a draft, and that’s the point

The preview is an editable text area. Fix the two lines that came out wrong, then copy or download. A conversion that claims to need no checking is either simplifying your document or guessing without saying so.

Straight answers

Will tables come out as tables?

No. A table arrives as its rows in reading order, one line per row, with columns separated by the spacing they had on the page. Rebuilding the pipes is manual work, and honest work, because a bad automatic table is worse than no table.

Why is there a heading at the top of every page?

That’s the page marker: two hashes and the page number. It makes your editor’s outline show the document’s shape, and it makes citations easy. It can be switched off if pages are an artifact of printing.

What happens to bold and italic text?

The words survive without markup. Nothing in a PDF reliably distinguishes emphasis from a font choice, and guessing turns a technical document into asterisks in the wrong places.

Can I get the images out too?

Not from here. The PDF-to-image tool turns pages into PNG or JPG if you need the visual, and Markdown can then point at those files.

Does it work on a scanned PDF?

Not on its own — a scan has no text layer, so there’s nothing to convert. Tick the OCR option and a recognition engine runs over those pages in your tab, and the Markdown export then writes them like any other page. It reads printed English, and a heading guessed out of a scan is rougher than one read from a real text layer.

More on extract text

The full tool, with every option: Extract Text.