Files & data
quickOCR
Pull text out of a screenshot without sending the screenshot anywhere.
- Runs at
- forhadkhan.github.io
- Cost
- Free · MIT licensed
- Your data
- Stays in your browser
Someone sends a screenshot of an error message instead of the text. You need the text. Every OCR site wants you to upload the image first.
quickOCR runs the recognition in your browser.
What it does
Pasting from the clipboard is the main path: screenshot, Ctrl+V, text. You can
also drag a file in or pick one. Zoom and pan the source image to check what it
read, edit the extracted text in place before copying, and go back to earlier
results through the scan history.
Nothing is uploaded
Tesseract.js is the Tesseract OCR engine compiled to WebAssembly, running in your tab on your CPU. The image is never transmitted.
This is the whole reason to use it over a hosted service. Screenshots are one of the most casually over-shared kinds of file there is. They routinely contain account numbers, internal URLs, customer records, private messages, whatever else happened to be on screen at the time. “Upload your image to extract the text” is a much larger ask than it sounds like.
One caveat, stated plainly: the language model files come from a CDN on first use. That’s a request for a static public data file and your image isn’t part of it, but it does mean the tool isn’t fully offline until those are cached.
Accuracy, before you rely on it
Tesseract is good. It is not a modern cloud OCR service, and the difference shows up in predictable places.
Screenshots of digital text come out excellent, which is the case I built it for. Clean scans of printed documents are good. Photographs of paper get variable, because angle and lighting and shadow all hurt. Handwriting is poor enough that you shouldn’t bother. Dense tables and multi-column layouts usually lose their structure even when every individual word is right.
If you need reliable OCR over photographed documents at any volume, Google Cloud Vision or AWS Textract are meaningfully better, at the cost of sending them your images. This trades some accuracy for that not happening.
Read the output against the source either way. The zoom and pan view exists for exactly that.
How it’s built
React and Vite for the interface, Tesseract.js for recognition, Framer Motion for transitions. Static files on GitHub Pages, where the “backend” is your CPU.
The interesting constraint was the WASM bundle. Engine plus language data is large, so the app loads the interface first and only pulls the engine on the first scan rather than on page load. Expect a pause the first time round.
Rough edges
First scan is slow, since the engine and language data have to download and initialise. Later scans in the same session are much faster. Large images take real time because it’s your CPU doing the work. English is the default and other languages need their data files loaded. There’s no PDF input, images only, so screenshot the page you need.
Verifying the privacy claim
Open devtools, watch the network tab, run a scan. You’ll see requests for the engine and the language data, and nothing at all carrying your image. Don’t take my word for it when checking takes ten seconds.
It’s slower than an online OCR service because their server has more CPU than your laptop does. For a screenshot that’s a second or two, which seems like a reasonable trade.