How OCR works
Optical character recognition (OCR) finds lines of text in an image, splits them into characters and matches the shapes against a trained language model. This tool uses the open-source Tesseract engine, running in your browser. The engine and the data for your chosen language are downloaded the first time and then cached.
Getting accurate results
- Use a sharp, well-lit image. Text should be at least about 20 pixels tall.
- Keep lines straight: crop and rotate skewed photos first.
- Dark text on a plain light background reads best.
- Choose the correct language, since each one has its own character set and dictionary.
What it can and cannot read
Printed and typed text, such as documents, book pages, screenshots, signs and receipts, usually reads well. Handwriting, decorative fonts and low-resolution images are much less reliable. Always proofread the result, especially numbers and names.
Private by design
The image is opened and processed by your browser on your own device. It is not uploaded to a server, so the tool also works offline once the page has loaded.
