English OCR

100% in your browser: nothing is uploaded

Loading…

The smallest and fastest pack on the site at 7.8 MB, trained on English alone. Pick it when you already know the page is English and want the shortest download and the quickest run; pick the Latin pack instead if the document might slip into another European language, and the Chinese pack if it might carry any CJK.

How it works

  1. Drop an image or PDF. Text detection runs immediately.
  2. Pick your language pack (downloaded once, 8–17 MB) and press Read text.
  3. Copy the lines, or download .txt / .json. PDFs are read cover to cover.

Questions

Why choose this over the default?
Download size and speed. The default reads Chinese, Japanese and English together and costs 16 MB; this one is 7.8 MB and has a smaller character set to search, so it is quicker per line. On English-only documents the accuracy difference is not what you are trading.
Does my file get uploaded?
No. The models and all processing run inside your browser on your own device. You can verify this yourself: open your browser’s DevTools, watch the Network tab, and process a file: there are no uploads. After the first visit most tools also work fully offline.
Which languages are supported?
106 languages across Latin, Cyrillic, Arabic, Devanagari, CJK, Greek, Thai, Tamil and Telugu scripts. Bengali is the one major gap, since no v5 model exists for it yet, and the picker says so rather than guessing.
How accurate is it?
On clean document text, the measured character error rate is under 2%, and every line carries a confidence score so you can see which ones to check. Handwriting and low-resolution photos will read worse.
Does it read scanned PDFs?
Yes. Digital PDFs are read from their own text layer, with exact characters at exact positions. Scanned pages go through detection and recognition, and mixed documents use both.

Related tools