How to make scanned PDFs and photos searchable, privately
A scanned PDF is a picture of a page, so search can't find anything in it. Text recognition (OCR) fixes that, and it can run on your own computer instead of someone's server.
You search a PDF for a word you can see on the page, and the search finds nothing. The PDF isn't broken: it's a scan. A scanner or a phone camera saves a picture of the page, and a picture has no words in it, only pixels shaped like letters.
The same goes for photos: a receipt you photographed, a whiteboard, a certificate on your phone.
Text recognition, explained in one minute
Text recognition, usually called OCR, looks at the picture and works out which letters the shapes are. The result is real text that can be searched, copied and quoted. Modern OCR engines use small neural networks trained on many fonts and languages, and on a clean scan they are very accurate.
The catch is where it runs. Many online converters ask you to upload the file, convert it on their server and send it back. For an ID card, a court decision or a bank letter, that's exactly the kind of document you shouldn't be uploading.
Running OCR in your browser
The open-source Tesseract engine, maintained for decades and now compiled to WebAssembly, runs inside a web page at close to native speed. AskFolder uses it that way: when it meets a scanned page or a photo, it draws the page in memory and reads it right there in your browser tab. The image never leaves your computer, and neither does the text.
Step by step
- Open askfolder.com/app.
- If your documents aren't in English, open Settings, then General, and choose the language of your scans. Czech, German, Spanish, French, Italian, Hungarian, Dutch, Polish, Portuguese and Romanian are built in, always alongside English.
- Drop in the folder with your scans and photos (PDF, JPG, PNG, WebP or BMP).
- Each scanned page is read in turn; the file list shows "read from a scan" or "read from the photo" as it goes.
- Ask in plain words. Answers from scans show their page like any other PDF.
Getting the best results
- Scan at 300 dpi if you can; phone photos work best taken straight on, in good light.
- Choose the right language: accented letters like "á", "ő" or "ß" are read far better with the language's own model.
- Very faint, skewed or handwritten pages may not be read. AskFolder says so on the file rather than skipping it silently.
Free and Pro
Free reads the first 3 scanned pages or photos in each folder, enough to see it working on your own documents. AskFolder Pro reads every one, in folders of any size. Nothing is uploaded either way.