Pathine

PDF to Text

Pull the text out of a PDF — including one that is only scans.

PDF Free · no signup Processed on our server Google Gemini
PDF to Text Read on your device

The text inside a PDF is read here, in this tab. Only a scanned page — one with no text in it at all — has to be sent anywhere, and the tool asks first.

Drop a PDF here

Contracts, statements, papers, manuals, scans. Any length: a PDF that carries its own text is read here from end to end, however long it is.

Reading the document…

That didn't work.

Extracted text

A PDF that contains its own text is read entirely in this tab and never leaves your device. Scanned pages hold no text to read, so those pages — and only those — are sent to a server to be recognised, and only when you ask for it. They are held in memory for the few seconds that takes, never written to disk, and never kept afterwards.

Keyboard shortcuts
/ or K
Search every tool
Run this tool, from anywhere on the page
S
Download the result
?
Open and close this list

About the Pathine PDF to Text

Pulls the text out of a PDF and hands it back as text you can copy, edit and search. Contracts, statements, papers, manuals, reports — anything where the words are trapped in a document that will not let you select them.

Nearly every PDF already contains its own text, and that copy is read here in this tab: no upload, no queue, no page limit, and the characters that come back are the exact characters the file was built with rather than a machine's best guess at them.

A scanned PDF is different, and the difference is the whole reason this tool has two halves. A scan holds no text at all — it is a photograph of a page — so there is nothing to read out of it and it has to be recognised instead. Those pages, and only those pages, are sent to a server to be read, after you press the button that says so.

Mixed documents are the normal case rather than the awkward one. A report with a signed page photographed and dropped back in is read both ways: the typed pages here, the photographed one by the model, and the result comes back in page order with nothing marked out of place.


When you'd reach for it

A contract or statement you need to quote

Getting the clause out exactly as written beats retyping it, and beats the transcription errors that come with retyping a reference number at the end of a long day.

A scanned document with no text in it

Old records, a posted letter, a form somebody photographed and mailed as a PDF. Search finds nothing in these because there is nothing in them to find until they are read.

Feeding a document to something else

A summariser, a search index, a spreadsheet, a script. Almost everything downstream wants plain text, and a PDF is the least convenient way to hand it over.

Counting or checking what is actually in a file

Word counts, a diff against last quarter's version, a search for every place a name appears. All of it needs the text out first.


How to extract text from a PDF

  1. Drop your PDF

    It opens on your device and the text comes out immediately. Nothing has been sent anywhere at this point, whatever the document turns out to contain.

  2. Read the summary

    It says how many pages had text in them and how many were scans. If every page had text, the tool is finished and your file never left the tab.

  3. Ask for the scanned pages, if there are any

    A button appears saying how many pages need reading and what that will cost from your daily allowance. Press it and those pages — not the whole document — are sent, read and slotted back into place. Leave it alone and nothing is sent.

  4. Copy or download

    Take the text to the clipboard, or save it as a .txt file. Nothing is kept here either way.


Common questions

Why does some of my PDF come out instantly and some slowly?

Because they are two different jobs. Pages with real text in them are parsed on your device, which takes milliseconds. Pages that are pictures have to be recognised by a model on a server, which takes a few seconds each. A document that is entirely scans is entirely the slow kind.

Does my PDF get uploaded?

Only if it contains scanned pages, and only after you press the button asking for them. Even then it is those pages that go — they are copied into a new PDF in your browser and that is what is sent, so a hundred-page report with two scans in it sends two pages. They are held in memory for the few seconds the reading takes, never written to disk and never kept.

Why is there a limit on the scanned pages?

Each batch of scanned pages is a call to a model, and those cost us money on a tool with no account and no payment. The allowance is per visitor per day, the tool says what is left, and the pages read on your device do not count against it at all — a nine-hundred-page PDF with a text layer is free and always will be.

Will the layout survive?

No, and nothing that claims otherwise is telling you the truth about what a PDF is. This gives you the words in reading order with the line breaks the file has, which is what you want for quoting, searching and counting. Columns, tables and figures are laid out by coordinates rather than by structure, so a two-column page comes back as two columns of prose in the order the file stores them. For a table specifically, Table to CSV is the right tool.

How accurate is it?

The pages with a text layer are exact — those characters are copied out of the file, not recognised. The scanned pages are very good and not infallible, and they are the ones worth a glance before you rely on them. Check numbers first: a misread digit in an amount is the error that actually costs something.

It says my PDF is password-protected.

Then it cannot be opened here, by either half of the tool. Open it in your usual PDF reader with the password, save a copy without one, and drop that in.

What about a photo or a screenshot rather than a PDF?

That is Image to Text, which takes JPG, PNG, WebP, HEIC and AVIF and does the same recognition on a single picture.

Related tools

Next door.