← All ZipSeek guides

PDF and OCR

How to Search Text in PDFs and Scanned Documents on Mac

Search native PDF text and locally OCR scanned pages, then move through matches on the exact page or image region.

Published Updated

Two PDFs can look identical and behave very differently in search. A digitally produced PDF usually contains selectable text. A scan may contain only page images, so there is no text for a normal content search to match. Mixed PDFs can contain both.

ZipSeek handles those cases in one search. It reads the native text layer where one exists and can use on-device OCR for scanned pages and embedded page images. The result records the page number and, when OCR supplies word coordinates, the matching region on the page.

Run one search across text and scans

Choose a PDF, an archive, or a folder of documents. Enter the phrase and keep Text in Images enabled when scans may be present. If you know every file has a real text layer, turning that option off skips OCR work while preserving ordinary PDF text search.

ZipSeek does not treat a missed text query as proof that every page needs OCR. Text pages use their native layer; pages without useful text, and text pages that contain relevant raster resources, are the OCR candidates. This keeps the expensive part of the search focused.

  • Use a short, distinctive phrase rather than a single common word when reviewing a large document set.
  • Leave PDF selected in the file-type filter; select Images as well when separate PNG, JPEG, HEIC, TIFF, BMP, GIF, or WebP files may contain the phrase.
  • Watch the bottom status area when OCR is active. It reports recognition progress separately from ordinary document parsing.

Review the exact page, not just a snippet

Selecting a PDF result opens it in the right-hand preview at the matching page. Native text matches are highlighted in the page text, while OCR matches use the recorded rectangles over the rendered page. The previous and next controls move through individual occurrences, not merely through lines that contain them.

Separate image files use the same idea: the preview scales the image to fit and maps the OCR rectangles onto that displayed size. This makes a result useful even when the recognized text is one label in a full-page scan.

OCR stays on this Mac

Recognition is performed with macOS Vision on the local machine. Documents, OCR text, and search terms are not sent to an external AI or document-processing service. Reusable OCR results are stored in the local search cache so the same unchanged image does not need to be recognized for every query.

Changing the source changes its content fingerprint, so stale OCR data is not reused. The cache is bounded and can be inspected, resized, or cleared in ZipSeek Settings.

Know where OCR can be wrong

OCR is an interpretation of pixels. Low resolution, skew, handwriting, faint photocopies, unusual fonts, and dense tables can all reduce accuracy. A missing result does not prove the words are absent from a poor scan, and a recognized result should still be checked against the page image.

If you only need to search one clean, text-based PDF, Preview already provides an excellent in-document search. ZipSeek becomes more useful when the same query must cover many PDFs, separate images, Office files, and archives while preserving each source path.