Skip to content
Pdfiqo

How to Make a Scanned PDF Searchable with OCR, Step by Step

Learn how OCR turns a scanned PDF into searchable, selectable text, how to get accurate results, and how to do it privately in your browser.

By Editorial Team, June 2, 2026. 8 min read.

A scanned PDF is really a stack of photographs of paper, so your computer sees pixels, not letters. To make it searchable you run optical character recognition (OCR), which detects the words in each page image and adds an invisible text layer on top of the picture. The page looks exactly the same, but now you can press Ctrl+F, select a sentence, or copy a paragraph into an email.

This guide explains what OCR actually does, how to prepare your scans so recognition is accurate, and how to run it without sending sensitive documents to a remote server.

How to tell whether your PDF needs OCR

Not every PDF that came out of a scanner needs OCR, and not every PDF that looks digital has real text. A quick test takes ten seconds:

  1. Open the PDF in any viewer.
  2. Try to click and drag across a line of text.
  3. Press Ctrl+F (Cmd+F on a Mac) and search for a word you can clearly see on the page.

If the selection draws a box around the whole page instead of highlighting words, or the search finds nothing, the page is an image. It needs OCR. If words highlight individually and search works, the file already has a text layer, and running OCR again will not help much.

Some documents are mixed. A contract might have typed pages exported from a word processor and a signed last page that was scanned. Check a few pages, not just the first. If only a few pages are scans, consider pulling those out with Extract pages, running OCR on them alone, and merging them back, because OCR converts every page it processes into an image with a text layer.

What OCR actually does to your file

It helps to understand the mechanics, because they explain both the strengths and the limits.

Step 1: Rendering each page as an image

The OCR engine needs a clean raster image of each page. It renders the page at a set resolution, measured in dots per inch (DPI). Higher resolution means more detail for the engine to work with, at the cost of speed and memory.

Step 2: Finding text regions and lines

The engine analyzes the layout: where the text blocks are, which direction lines run, and where one column ends and the next begins. Tables, sidebars, and captions make this harder.

Step 3: Recognizing characters and words

Modern engines recognize whole lines using trained models rather than matching one letter at a time. A language model then helps pick plausible words, which is why selecting the correct language matters so much. An engine set to English will struggle with German umlauts or Polish diacritics, and it will be close to useless on Japanese.

Step 4: Writing an invisible text layer

Finally, the recognized words are placed on the page as invisible text, positioned over the matching pixels. This is what people mean by a "searchable PDF" or "sandwich PDF." The visible layer is still a picture of the page, so the document looks like the original scan while the hidden text makes it searchable.

Preparing your scans for accurate recognition

OCR quality depends more on the input than on the software. A few habits make a noticeable difference.

Scan at a sensible resolution

For ordinary printed text, 300 DPI is the usual sweet spot. Below 200 DPI, small type and punctuation start to blur together. Going far above 300 DPI rarely improves accuracy for normal documents and makes files much larger.

Keep pages straight and well lit

Skewed pages, curled book spines, and shadows from a phone camera all reduce accuracy. If you are photographing documents, lay them flat, shoot from directly above, and use even light. If a page came out sideways or upside down, fix the orientation before OCR with our free Rotate PDF tool, since recognition works best on upright text.

Remove what you do not need

Blank separator sheets, cover pages, and duplicate scans waste processing time. Clean them out with Delete pages first. If the pages were fed through the scanner in the wrong order, Reverse PDF or Organize PDF can put them back in sequence.

Know what OCR cannot read well

Be realistic about certain content:

  • Handwriting. Standard OCR engines are trained on printed type. Neat block capitals may partly work, but cursive notes usually come out as noise.
  • Low-contrast or faded text. Carbon copies, thermal receipts, and old photocopies lose detail the engine needs.
  • Decorative fonts and stylized logos. These are often skipped or misread.
  • Complex tables. The words are usually recognized, but the reading order across cells may not match what you expect when you copy text out.

Running OCR on your own computer

There are several ways to add a text layer, and it is worth knowing the options.

Built-in and desktop options

  • macOS can recognize text in images through Live Text in Preview and Photos, which is handy for copying a line or two. It does not always produce a full searchable PDF, though.
  • Full desktop PDF editors usually include an OCR feature that writes a text layer into the file. These work well but are often part of a paid subscription.
  • Many scanners ship with software that can save straight to a searchable PDF. If you scan often, check your scanner settings before anything else.
  • Open-source command-line tools exist for technical users who want to batch-process large archives.

In your browser, without uploading

Our free OCR PDF tool runs the recognition engine directly in your browser. The file is read from your disk into the page, processed on your own device, and saved back as a new PDF. It is never uploaded to a server, which matters when the scan is a medical record, a lease, a bank statement, or anything with an ID number on it.

Here is the process:

  1. Open the OCR PDF tool and choose your scanned PDF.
  2. Pick the document language. The tool supports 13 languages: English, German, French, Spanish, Italian, Portuguese, Polish, Russian, Ukrainian, Dutch, Turkish, Japanese, and Simplified Chinese.
  3. Choose the resolution. 300 DPI is the default and the better choice for small print. 200 DPI is faster and works fine for large, clean text.
  4. Start the process and wait while each page is recognized. Your browser tab does the work, so a long document on an older laptop will take a while.
  5. Download the searchable PDF and test it with Ctrl+F.

The first run may take a little longer because the browser has to load the recognition engine and the language data. After that, work happens locally.

Checking and using the result

Once you have the searchable file, spend a minute verifying it.

  • Search for a few distinctive words such as a name, an invoice number, or a date. If they turn up, the text layer is working.
  • Copy a paragraph into a plain text editor. This shows you exactly what the engine recognized, including any mistakes like "rn" read as "m" or "0" read as "O."
  • Spot-check numbers. OCR errors in prose are usually easy to notice, while an error in an account number is not. If exact figures matter, compare them against the image by eye.

Extracting the text

If what you really want is the words themselves, run the searchable PDF through our free PDF to Text tool to get a plain .txt file, page by page. For an editable document with paragraphs and headings, PDF to Word is the next step. Keep in mind that both tools depend on the text layer, so a scan that has not been through OCR will come out empty.

Keeping the file size manageable

Scanned PDFs are heavy because every page is an image. The OCR tool renders each page fresh at the resolution you chose and pairs it with the text layer, so the output can end up larger than the original, particularly at 300 DPI. If the file is too big to email, our Compress PDF tool re-encodes the page images to shrink it. Try the lighter setting first and check that small print is still readable before sending.

Common problems and how to fix them

Search finds nothing after OCR. Make sure you opened the new file, not the original. Some viewers also cache documents, so close and reopen it.

The text is gibberish. The language setting was probably wrong, or the page was rotated. Fix the orientation and rerun with the right language.

Only part of the page was recognized. Very light text, colored backgrounds, or text over photos may be missed. Rescanning at higher contrast often helps more than any software setting.

Processing is slow or the tab crashes. OCR is demanding, and browser-based processing is limited by your device's memory. Split a very long document into smaller parts with Split PDF, process each one, then combine them again with Merge PDF. Closing other heavy tabs also helps.

The document mixes languages. Choose the language that makes up most of the text. Words in the other language will still often come through, especially if they share an alphabet, but expect more errors there.

Why the privacy angle matters for scans

People tend to scan exactly the documents they would never post publicly: passports for a visa application, tax forms, medical referrals, signed agreements. Many OCR services work by uploading the file to a server, processing it there, and sending the result back. That may be fine for a restaurant menu, but it means a copy of your document exists, at least temporarily, on someone else's infrastructure.

When recognition runs in the browser, that trip does not happen. You can confirm it yourself by opening your browser's developer tools, switching to the Network tab, and watching for requests while the tool processes your file. You will see the engine files load, but not your document going out.

Key takeaways

  • A scanned PDF is a set of images, and OCR adds an invisible text layer so you can search, select, and copy.
  • Accuracy depends mostly on the scan: 300 DPI, straight pages, good contrast, and the correct language.
  • Handwriting, faded copies, and complex tables remain difficult for any OCR engine.
  • Always verify important numbers and names against the page image.
  • In-browser OCR keeps sensitive scans on your device, with speed limited by your computer's resources.

When your next stack of paperwork comes out of the scanner, run it through OCR PDF before you file it away. Months later, when you need to find one clause or one invoice number, a quick search will get you there.