PDF OCR
Turn a scanned PDF into a searchable, selectable one. The OCR runs locally in your browser, so the file never leaves your device.
Nothing is uploaded
The PDF OCR engine runs as WebAssembly inside your browser. Your scan is never sent to a server.
See what it read
Every recognized word is highlighted on the page, so you can check the result before downloading.
11 languages
Latin, Cyrillic and CJK scripts are all supported, including Chinese, Japanese, Korean and Russian.
Free, no signup
No account, no watermark, no page cap. The only real limit is your own device memory.
In short: PDF OCR reads the text off a scanned page and writes it back into the PDF as an invisible text layer. The page still looks exactly the same, but you can finally search, select and copy it. This tool does the whole job inside your browser.
Why Scanned PDFs Break Your Workflow
You get a contract back from a client. You need to find one clause. You hit Ctrl+F, type the word, and the PDF tells you there are no results — even though you can see the word right there on the screen.
That is the moment most people go looking for PDF OCR. The document is not broken. It simply never contained any text to begin with.
A scan is a picture, not a document
When a scanner or a phone camera captures a page, it records pixels. The letters you see are shapes, not characters. To the PDF reader, the whole page is one flat image.
What you lose without text recognition
No search. No copy and paste. No text selection for quoting. Screen readers announce nothing. Your document management system indexes an empty file. Every one of those features depends on the page containing real characters.
Where it hurts most
Signed contracts, paper invoices, receipts, medical records, academic papers, old archives. These are exactly the documents you most need to search later, and exactly the ones that arrive as scans.
What Is PDF OCR?
PDF OCR is the process of reading text out of a scanned page image and storing it back into the PDF as a real, invisible text layer — so the page still looks identical, but you can now search, select and copy the words on it.
OCR stands for optical character recognition. The engine looks at the shapes on the page, matches them against the letterforms of a language, and outputs characters plus the exact position of each word.
How the invisible text layer works
The recognized words are drawn back onto the page in PDF rendering mode 3, which means "draw this text but do not show it". Each word sits exactly on top of its picture in the scan.
That is why a properly OCR-processed file looks completely unchanged, yet highlights correctly when you search it. You are selecting the invisible layer, not the image.
Searchable PDF vs plain text extraction
When you want a searchable PDF
You need to keep the original look — signatures, stamps, letterhead, layout. Anything that has to stay presentable or legally recognizable should stay a PDF.
When you just want the raw text
You are pasting into an email, a spreadsheet or a note. Then skip the PDF and take the plain text export instead. This tool gives you both from the same run.
What OCR cannot fix
OCR reads what is there. It cannot recover text that the scan blurred away, and it does not proofread. A crooked, low-resolution photo of a page will produce crooked, low-quality results.
How to OCR a PDF in 3 Steps
The whole process happens on your own machine. Nothing is queued, nothing is uploaded, and there is no waiting room.
- 1
Open your scanned PDF
Drag the file onto the box above, or pick it from your device. It loads straight into the browser tab — there is no upload step at all.
- 2
Pick the document language
Choose the language actually printed on the page, not your interface language. The engine downloads that language pack once and keeps it cached for next time.
- 3
Run OCR and download
Watch it work page by page, check the highlighted words, then take the searchable PDF or the plain text. Both are generated locally.
Pro tip — Language choice matters more than most people expect. Running an English pass over a German page will mangle every umlaut. If your results look like nonsense, the language is usually the reason.
Note — The first run on a new language downloads a language pack of under 3 MB. After that it is instant, and it works offline.
See What the OCR Actually Read
Most PDF OCR tools hand you a file and wish you luck. You only find out whether it worked after downloading, opening the PDF and searching it yourself.
That is a bad deal, because scan quality varies enormously and PDF OCR fails quietly. A tool that silently returns 60% accuracy looks exactly like one that returns 99%.
The text layer, made visible
Turn on "Show text layer" and every recognized word is boxed directly on the page. What you see is precisely what got written into the PDF.
How to spot a bad scan before you download
Words the engine is unsure about are marked in red. Hover any box to see what it actually read and how confident it was.
A page with a few scattered red boxes is fine. A page that is mostly red means something is wrong — usually the wrong language, or a scan that needs to be redone at a higher resolution.
What You Can Do With a Searchable PDF
Once the text layer exists, the document starts behaving like a real document everywhere you use it.
Search inside the document
Ctrl+F works in any PDF reader, on any device. Finding one clause in a 60-page scanned contract stops being a manual scroll.
Copy and quote without retyping
Select a paragraph and paste it into an email or a report. No transcription, no typos introduced along the way.
Make scans accessible
Screen readers can only read a text layer. Running PDF OCR is often the single biggest accessibility improvement you can make to an archive of scans.
Feed clean text into other tools
Search indexes, note apps, translation tools and spreadsheets all need characters, not pixels. The plain text export drops straight into them.
Your scan never leaves this tab
Scanned documents are usually the most sensitive files people own: signed agreements, passports, tax paperwork, medical letters. Uploading those to an unknown server to make them searchable is a strange trade to make.
This tool loads the PDF OCR engine into your browser as WebAssembly and runs it there. The language packs come from our server, but your document does not go anywhere. You can disconnect from the network after the page loads and it will still finish the job.
Browser-Based PDF OCR vs Upload-Based OCR Tools
Local processing is the right default for private documents, but it is not free of tradeoffs. Here is the honest comparison.
| Aspect | This tool (in browser) | Upload-based OCR tools |
|---|---|---|
| Where your file goes | Never leaves your device | Uploaded to a remote server |
| Works offline | Yes, after the first load | No |
| File size limit | Bounded by your device memory | Bounded by the service's plan |
| Account required | No | Often, for larger files |
| Speed on very large PDFs | Slower — it uses your CPU | Faster — it uses their servers |
| Accuracy on difficult scans | Good on clean print | Often better — commercial engines |
The last two rows are real. A commercial OCR service with a tuned engine will beat an open-source one on a creased fax from 1998, and a server farm will always out-run a laptop on a 500-page file. If your document is not sensitive and accuracy on a bad scan matters more than privacy, use one of those instead.
How to Get the Best OCR Accuracy
PDF OCR quality is decided mostly before you press the button. These four things account for most of the difference.
- 1
Scan at 300 DPI or higher
This is the single biggest factor. At 150 DPI, character edges start merging; at 72 DPI, most engines guess. 300 DPI is the standard target for text recognition.
- 2
Straighten and clean the page first
A page rotated even two degrees measurably lowers accuracy. Turn on deskew for crooked scans, and background removal for grey or speckled ones.
- 3
Use one language per pass
Mixed-language packs dilute accuracy. If a document is mostly English with a few French names, run it as English.
- 4
Check the text layer, then re-run
Look at the highlighted words before you download. If a page is full of red boxes, change the language or the scan settings and run it again — it costs you nothing.
Complete PDF Tool Suite
Discover our comprehensive collection of PDF tools designed to handle all your document needs
PNG to PDF
Bind PNG images into a single, print-ready PDF
JPG to PDF
Convert JPG images to PDF format
Merge PDF
Combine multiple PDF files into one
Compress PDF
Reduce PDF file size efficiently
PDF to PNG
Convert PDF pages to PNG images
PDF to JPG
Convert PDF pages to JPG images
PDF to Text
Extract text content from PDF files
Split PDF
Split PDF into separate pages
Edit PDF
Edit and annotate PDF documents
Organize PDF
Organize and rearrange PDF pages
Rotate PDF
Rotate PDF pages and save permanently
Page Numbers
Add page numbers to PDF with live preview
Watermark PDF
Add a text watermark to PDF with live preview
HEIC to JPG
Convert iPhone HEIC photos to JPG
Delete PDF Pages
Remove unwanted pages and download a clean PDF
Extract PDF Pages
Save selected pages as a new PDF or separate files
Sign PDF
Draw, type, or upload a signature and place it on any page
Resize PDF
Change page size to A4, Letter, or custom dimensions
Crop PDF
Trim margins, drag-select, or auto-crop white space
Flatten PDF
Make forms read-only — keep text searchable or lock to image
PDF Metadata
View, edit, or strip author, title, dates and other metadata
Grayscale PDF
Convert to grayscale or black & white to save colour ink
Extract Images from PDF
Pull embedded photos out of a PDF and save as PNG or JPG
WebP to PDF
Convert WebP images to a PDF, merge many into one file
Protect PDF
Add an open password to your PDF, entirely in your browser
Unlock PDF
Remove a known password or restrictions from a PDF, in your browser
HEIC to PDF
Convert iPhone HEIC photos to PDF
AVIF to PDF
Convert AVIF images to PDF
TIFF to PDF
Convert TIFF images and scans to PDF
BMP to PDF
Convert BMP bitmap images to PDF
OCR PDF
Make a scanned PDF searchable, entirely in your browser
Redact PDF
Permanently remove sensitive text from a PDF, in your browser
Compress Image
Shrink JPEG, PNG and WebP in batch, entirely in your browser
Resize Image
Change image dimensions in batch, with presets, in your browser
Frequently Asked Questions
Everything people ask before running PDF OCR on a scan.
How do you OCR a PDF for free?
Open the file above, pick the document language, and press Run OCR. The whole process happens in this browser tab and costs nothing. You get back a searchable PDF and a plain text file, with no account and no page limit.
Is this PDF OCR tool really free? What is the limit?
It is free with no page cap and no watermark. We do not impose a quota because we do not carry the cost — your device does the work. The practical limit is memory: very large scans on an older phone can run out of it, in which case run a page range instead of the whole file.
Do my files get uploaded to a server?
No. Your PDF never leaves your browser. The OCR engine is WebAssembly running locally. The only things downloaded from us are the engine and the language pack. You can verify this by disconnecting from the network once the page has loaded — OCR still works.
How accurate is the PDF OCR?
Very good on clean printed text, weaker on handwriting and poor scans. A 300 DPI scan of ordinary printed text typically recognizes in the nineties. Handwriting, faded fax paper, heavy tables and low-resolution photos degrade quickly. Use the text layer view to judge your own document rather than trusting an average.
Can I convert a scanned PDF to Word with OCR?
Not directly — this tool outputs a searchable PDF and plain text, not a .docx. For most purposes the plain text export is what you actually want, and you can paste it straight into a document. If you need the raw text on its own, our PDF to text tool covers that case.
Which languages does the PDF OCR support?
Eleven, covering Latin, Cyrillic and CJK scripts. English, Simplified and Traditional Chinese, Spanish, Portuguese, German, French, Italian, Japanese, Korean and Russian. Non-Latin scripts get a full text layer, so Chinese and Russian scans are genuinely searchable, not just partially.
Why is the first run slower than the second?
The first run downloads the OCR engine and the language pack. After that both are cached in your browser, so later runs start immediately and work without a network connection. Switching to a new language downloads that pack once.
What should I do after I OCR a PDF?
Usually compress it, or pull the text out. The text layer only adds a few kilobytes per page, but the underlying scan is often large. Compressing afterwards keeps the searchability and cuts the file size.
Related Tools
Other things you might want to do with the same document.
- PDF to Text — pull the raw text out of a PDF that already has a text layer.
- Compress PDF — shrink the scan after OCR without losing the text layer.
- PDF to PNG — export pages as images if you want to clean them up before scanning again.
- Grayscale PDF — strip colour from a noisy scan, which often helps recognition.
- Merge PDF — combine several scanned pages into one document before running OCR.
- Redact PDF — permanently remove sensitive text once the scan is searchable.
Make your scans searchable
Run PDF OCR right here in your browser. No upload, no account, no page limit.
Open the PDF OCR tool