PDF to Text Converter
This PDF to text converter works in one click: drop in a PDF, click “Extract text”, done. Broken lines are joined into clean paragraphs, and scanned pages are read by text recognition (OCR). Everything runs in your browser – the file is never uploaded.
Source: Adobe – ISO 32000-1:2008 Document management – Portable document format (PDF 1.7). Updated: .
How it is calculated
Copying text out of a PDF often ends in a mess: hard line breaks mid-sentence, hyphenated words (“docu-ment”), mixed-up columns. This tool reads the PDF’s text layer with pdf.js, groups the text pieces into lines by their position on the page and joins them into paragraphs. You can edit the result right away, copy it or save it as a TXT file (UTF-8).
How it works
- Drop in or choose a PDF.
- Pick a text layout: “Paragraphs” gives running text you can keep writing, “Lines as in the PDF” keeps every break – good for addresses, poems or lists.
- Optionally tick “Mark page breaks” to put “--- Page 3 ---” before each page.
- Click “Extract text”, check the result, copy or save it. You can change the options afterwards without reading the PDF again.
Scanned PDFs: text recognition (OCR)
A scan is a picture without a text layer. If a page has fewer than 20 characters, the tool renders it at about 300 dpi and lets Tesseract read it – the same free OCR engine used by Image to text. Choose the document’s language for it. Recognition takes a few seconds per page; it is very accurate on clean scans, not on handwriting or poor copies.
Honest limits
- Tables and multi-column layouts are read line by line, so columns can run into each other. For editable documents with headings, use PDF to DOCX.
- Some PDFs use fonts without a character map, producing boxes or gibberish. Tick “Read all pages with OCR” in that case.
- Password-protected PDFs open with their password; the tool does not bypass copy protection.
Privacy
Your PDF stays on your device. Only the PDF engine and, if needed, the text recognition are loaded from this server – no third-party service sees your document. For structured text for notes or AI chats, use PDF to Markdown.
Frequently asked questions
How do I convert a PDF to text for free?
Drop your PDF in here and click “Extract text”. Copy the text with one click or save it as a TXT file. No sign-up, no upload.
How do I convert a scanned PDF to text?
Keep “Read scanned pages with text recognition (OCR)” switched on and choose the document’s language. Pages without a text layer are then read by OCR automatically.
Why does copied PDF text have so many line breaks?
PDFs store text line by line, not as paragraphs. With the “Paragraphs” layout the tool joins lines that belong together and removes end-of-line hyphenation.
Can I paste the PDF text into Word or Google Docs?
Yes: click “Copy” and paste. If headings and paragraphs should survive as a Word file, use “PDF to DOCX”.
Is my PDF uploaded?
No. Reading and text recognition run entirely in your browser.
Why do I only see boxes or strange characters?
The PDF uses fonts that are not mapped to real letters. Tick “Read all pages with OCR” and extract again – text recognition then reads what is printed instead of the broken text layer.
Sources and legal basis
- Adobe – ISO 32000-1:2008 Document management – Portable document format (PDF 1.7)
- pdf.js – getTextContent / getAnnotations (Mozilla, Apache-2.0)
- pdf.js 6.3.289 licence (Apache-2.0)
- Tesseract.js – OCR library for the browser (GitHub, Apache-2.0)
- Tesseract OCR documentation – Improving the quality of the output
- Unicode Standard – UTF-8 encoding form (TXT export)
- Open-source licenses of the libraries used (MIT, LGPL-3.0)
As of:
Related tools
- PDF to Markdown ConverterConvert PDF to Markdown: headings, paragraphs, lists and links as .md – for Obsidian, GitHub, notes or AI chats like ChatGPT and Claude. OCR for scans.
- Convert PDF to DOCXTurn a PDF into an editable DOCX file for Word, LibreOffice or Google Docs – right in your browser, scanned pages via OCR. No upload, no sign-up, free.
- Image to text converterConvert an image to text: copy text from photos, scans and screenshots. OCR for English, German, French, Spanish and Portuguese – in your browser, no upload.
- PDF to JPG ConverterConvert PDF to JPG or PNG: every page as its own image, 72 to 300 dpi, one by one or as a ZIP. Free and right in your browser – nothing is uploaded.
- Text CleanerRemove line breaks, extra spaces, empty lines and invisible characters, delete duplicate lines and sort lists. Free, instant and private in your browser.
- Add Border to ImageAdd border to image online: white or colored border, instant-photo style, rounded corners and a transparent frame, with preview. Free, no upload.