Extract text
Get the words out of a PDF and into something you can edit. Page by page, in reading order, as plain text or Markdown.
Drop the PDF you want the text from
or choose a file
Reports, papers, e-books, invoices. Scanned pages have no text layer; tick "OCR scanned pages" below to read them.
—
—ℹ️
What "reading order" means here. A PDF has no paragraphs and no columns, just text runs parked at coordinates. Extraction means guessing the order back from those coordinates: group what shares a baseline, sort left to right, and it usually comes out right. Two-column layouts, tables, and captions mixed with body text are where the guess shows. Nothing is stitched into paragraphs, and a wrong guess there costs you a heading swallowed by the paragraph above it, which is worse than an honest line break.
The header/footer filter only considers the first three and last three lines of each page, and only drops a line if it repeats on at least 60% of them. Heading inference compares against the median body size, so it will occasionally promote a bold caption. Both are reported in the result so you can see what was touched.
OCR runs the ~7 MB recognition engine in your tab, downloaded on first use only. It reads printed English text; handwriting and low-quality scans will come out imperfect, so check the result.
Editable: fix anything the layout guess got wrong before you export or copy.
