Extract text from a PDF

Pull the text out of a PDF. Runs on your device.

How it works

  1. Add your PDF
  2. Choose pages
  3. Copy or download

Which PDFs work

This tool reads the text that is stored inside a PDF. Most PDFs saved from a word processor, a web page, a spreadsheet or an e-book have that text layer, and for those the tool returns the words as they are stored, one page after another. It does not recognise letters in pictures. A PDF made by scanning paper, or by taking photos of pages, holds a picture of each page and no text, so there is nothing to extract. The tool tells you which pages look like that, for example page 4 of a 12-page file, and still reads the others. It also cannot open a PDF that is protected with a password; remove the password first. One PDF is read at a time, and a file over 200 MB or over 1000 pages asks you to confirm before it loads.

How lines are rebuilt

A PDF does not store lines or paragraphs. It stores small pieces of text, each with a position on the page. The tool groups pieces whose heights are close into one line, sorts each line from left to right, and puts a space where the gap between two pieces is wider than a small fraction of the text size. A much wider gap, such as the space between a label and a number in a table, becomes three spaces. When the distance between two lines is clearly larger than usual, it adds a blank line, which usually marks a new paragraph. You can switch blank lines off, and you can add a line such as --- Page 3 --- before each page. Rotated pages are turned back so the text reads in its normal direction.

Where the order can be wrong

The tool reads across the page, line by line. On a page with two columns, a line of the left column is followed on the same row by the matching line of the right column, so sentences are cut in the middle and joined with the other column. Tables come out as rows with wide gaps, not as cells. Text in right-to-left scripts such as Arabic or Hebrew is not supported and may come out in the wrong order. Text drawn sideways on a page that is otherwise upright, such as a margin label, can end up out of place. Unusual fonts and ligatures can produce odd characters, and the tool shows them as they were extracted instead of cleaning them up. Always read the result once before you rely on it.

Scanned PDFs and privacy

If a page has no text layer, the tool lists it as probably scanned. To get text from a scan you need character recognition, which works on pictures. Turn the pages into images first, then use a tool that reads text from images. The PDF is opened and read by code running in this tab. There is no upload step and no server that receives your file or your text. The page counts visits with Google Analytics, using rough groups such as 1 to 10 MB, never a file name or any text from the document. The text is kept in memory while the page is open, and it is not stored or put in the address bar.

Frequently asked questions

Why is the text empty?

The PDF probably holds pictures of pages, such as a scan or a phone photo saved as PDF, and has no text layer. The tool lists the pages it could not read, for example page 4. Text from a picture needs character recognition, which this tool does not do.

Why are the columns jumbled?

The tool reads each row across the whole page. On a two-column page it joins a line of the left column with the matching line of the right one, so a sentence is split in the middle. If the order matters, copy one column at a time from the original PDF.

Can I extract only some pages?

Yes. Type pages such as 1-3, 5, 8-10 in the Pages box. A range of 2-3 from a 12-page file returns only those two pages, and with the page lines switched on they are labelled --- Page 2 --- and --- Page 3 ---, matching the PDF. Leave the box blank to read every page.

How is this different from copy and paste?

Selecting text in a PDF viewer often breaks lines in odd places or skips pages you did not scroll to. This tool reads a whole page range at once, rebuilds the lines from the position of every piece of text, and gives you a .txt file in UTF-8. It does not fix the order on multi-column pages.

What can I do with a scanned PDF?

Turn its pages into JPG images, then run a tool that recognises text in images. Recognition can make mistakes, so check names and numbers against the page.

Is my PDF uploaded?

No. The PDF is read in your browser tab, and the text is made there. The site counts visits with Google Analytics using rough size groups, never a file name or any of the text. A very long result shows only its first 200,000 characters on screen, but Copy and Download use all of it.