Extract Text from PDF
Pull the text out of a PDF with its line breaks and paragraphs intact, and read scanned pages in Korean or English with built-in OCR. Free, no upload.
Drop a PDF here or click to browse
🔒 Processed in your browser — never uploaded. Text recognition runs on your device too.
Get the text out of a PDF — laid out the way the page was
Drop in a PDF and this tool extracts all of its selectable text, page by page, so you can copy it or
download it as a .txt file. Everything runs in your browser — your file is never uploaded.
Scanned PDFs
If a PDF is made of scanned images with no text in it, the tool notices, says so, and offers to read the letters out of the pictures instead. Pick the language of the document — 한국어, 한국어 + English, or English — and press Read the text from the images. Korean is handled by PaddleOCR's PP-OCRv5 models, English by Tesseract; both are downloaded once and then kept by your browser, and both run on your own device, so a confidential scan stays on your computer just like an ordinary PDF does.
Only the pages that carry no words of their own are read. A report with three scanned pages inside it keeps the exact text of every other page, because the text a PDF holds is what was typed and recognition is a guess at it. If you do not trust the text a PDF carries — a broken text layer is a real thing — tick Read every page, even the ones that already have text and the guess wins instead.
A searchable copy of the scan
Beside that button is Download a searchable PDF. It gives back the same document, looking exactly as it does now, with the recognised words written underneath the picture where they belong — so Ctrl+F finds them and selecting a line copies real characters. Nothing visible changes and nothing is uploaded; the recognition runs on your device, which is the part every other site charges an upload for.
A PDF with a password
If the file is locked, the tool asks for the password instead of calling the file unreadable. It is used in your browser, like everything else here, and goes nowhere. A locked file cannot be saved as a searchable copy — take the password off it first with the PDF password tool — but its text reads out normally.
Frequently asked questions
Can I save the text as a file?
Yes — besides copying, you can download the extracted text as a .txt file, which is handy for long documents.
Why did it find no text?
The PDF is probably a scan — just images of pages with no real text layer. When that happens the tool says so and offers to read the words out of the page images instead, so you still get your text; it simply takes longer than pulling out text that was already there.
Does it work on Korean scans?
Yes. Choose 한국어 (or 한국어 + English for a mixed document) and the pages are read with PaddleOCR's PP-OCRv5 models, which are built for Hangul. The models are about a 32 MB download the first time and are kept by your browser afterwards; after that, expect a few seconds per A4 page. Nothing is uploaded — the recognition runs on your own device.
Can I get a searchable PDF back, not just the text?
Yes. When the document is a scan, "Download a searchable PDF" gives back the same file, looking exactly as it does now, with the recognised words written invisibly underneath the picture so a search finds them and a selection copies real letters. It reads the pages again at its own resolution, so it is a second pass rather than a re-use of the text on screen.
Does it read pages that already have text?
Not unless you ask. By default only the pages with no words of their own go through recognition, so a report with a few scanned pages inside it keeps the exact text of all the rest. Tick "Read every page, even the ones that already have text" when the text layer itself is the thing you do not trust.
What about a password-protected PDF?
It asks for the password and opens the file with it. The password is used in your browser and is not sent anywhere. The one thing it cannot do with a locked file is save a searchable copy of it — remove the password first with the PDF password tool.
Which languages can it recognise?
Korean and English. Every language is a recognition model your browser has to download, so this is not a list that can be as long as a server-side service's: the two here are about 32 MB and 10 MB, and they are served from this site so nothing goes to anyone else. The Korean model also reads the Latin letters and numbers mixed into a Korean page, which is why 한국어 and 한국어 + English use the same one.