Qovox OCR SDK
OCR and document data extraction that runs in the browser. Get a license key from your dashboard, add your domain, and you are ready.
How it works
Qovox OCR is a JavaScript library, not a cloud service. Your page loads it, and the work happens in the visitor's browser. The only server involved is a small license API that never sees a document.
- Start.
QovoxOCR.init({ licenseKey })sends the key and the page's origin toPOST /verify-license. The answer says whether the key is valid for this domain and returns the plan limits. If our server cannot be reached, scanning keeps working (fail open). Only an explicit "no" (unknown key, wrong domain, revoked, quota used) stops it. - Prepare. The image or PDF is decoded locally. Large images are shrunk to 2000 px, converted to grayscale and contrast-stretched in a Web Worker, so the page does not freeze.
- Read. A WebAssembly OCR engine runs in its own Worker. With
languages: 'auto'it starts with Arabic and English and only loads heavier language models when the result looks wrong. Models are cached in IndexedDB after the first download. - Structure. Text blocks with positions and confidence are turned into
text,markdownandjson, and invoices, ID cards and passports (MRZ with check digits) are parsed into fields. - Count. After a scan the SDK sends
POST /heartbeatwith a page count, so quotas and your dashboard stay accurate.
| Leaves the device | Never leaves the device |
|---|---|
| License key, page origin, SDK version, number of pages scanned | Images, PDFs, extracted text, fields, scanner pages, generated PDFs |
| One-time downloads of the OCR engine and language models from a public CDN (cached afterwards) |
Install
Script tag
<script src="https://qovox-ocr.web.app/qovox-ocr-sdk.js"></script>
npm
npm i qovox-ocr
import QovoxOCR from 'qovox-ocr'; // ESM / bundlers
const QovoxOCR = require('qovox-ocr'); // CommonJS
Quick start
<input type="file" id="f" accept="image/*,application/pdf">
<script>
const ocr = await QovoxOCR.init({ licenseKey: 'QVX-XXXX-XXXX-XXXX-XXXX' });
document.getElementById('f').addEventListener('change', async (e) => {
const result = await ocr.scan(e.target.files[0], {
onProgress: ({ status, progress }) => console.log(status, Math.round(progress * 100) + '%')
});
console.log(result.text); // plain text
console.log(result.markdown); // Markdown
console.log(result.json); // structured data, see below
});
</script>The first scan downloads the OCR engine and language data (about 10 MB) and the browser caches them. The document itself is never uploaded.
The result
{
text: "…",
markdown: "## INVOICE\n**Invoice No:** INV-2026-0042 …",
json: {
documentType: "invoice", // invoice | id | passport | document
fields: { invoiceNumber, issueDate, dueDate, currency, subtotal, tax, total, lineItems: [...] },
confidence: 93, // 0-100, mean over pages
language: "eng", pageCount: 1, pagesProcessed: 1,
blocks: [{ text, confidence, bbox: { x0, y0, x1, y1 } }],
pages: [{ page, width, height, confidence }]
},
confidence: 93, ms: 1840,
privacy: { documentBytesUploaded: 0, processedLocally: true }
}
Fields by type: passport (from the MRZ): surname, givenNames, passportNumber, nationality, dateOfBirth, sex, expiryDate. id: fullName, idNumber, dateOfBirth, expiryDate, nationality, sex. invoice: vendor, invoiceNumber, issueDate, dueDate, currency, subtotal, tax, total, lineItems. Only fields that were found are present, so always check for undefined and use confidence to route weak scans to a person. Force a type with scan(file, { documentType: 'invoice' }).
React
import { useEffect, useRef, useState } from 'react';
import QovoxOCR from 'qovox-ocr';
export function Scanner() {
const ocr = useRef(null);
const [out, setOut] = useState(null);
const [busy, setBusy] = useState(false);
useEffect(() => {
let alive = true;
QovoxOCR.init({ licenseKey: import.meta.env.VITE_QOVOX_KEY })
.then((o) => { if (alive) ocr.current = o; else o.destroy(); });
return () => { alive = false; ocr.current && ocr.current.destroy(); };
}, []);
async function onFile(e) {
setBusy(true);
try { setOut(await ocr.current.scan(e.target.files[0])); }
finally { setBusy(false); }
}
return (<>
<input type="file" accept="image/*,application/pdf" onChange={onFile} disabled={busy} />
{out && <pre>{out.markdown}</pre>}
</>);
}Vue 3
<script setup>
import { ref, onMounted, onBeforeUnmount } from 'vue';
import QovoxOCR from 'qovox-ocr';
let ocr; const result = ref(null), busy = ref(false);
onMounted(async () => { ocr = await QovoxOCR.init({ licenseKey: import.meta.env.VITE_QOVOX_KEY }); });
onBeforeUnmount(() => ocr && ocr.destroy());
async function onFile(e) { busy.value = true; try { result.value = await ocr.scan(e.target.files[0]); } finally { busy.value = false; } }
</script>
<template>
<input type="file" accept="image/*,application/pdf" @change="onFile" :disabled="busy" />
<pre v-if="result">{{ result.markdown }}</pre>
</template>Document scanner
Turn a camera frame or a phone photo of a paper document into a straight, clean page, then into a multi-page PDF or text. It uses plain canvas code (no extra download, no eval), so it works under strict Content Security Policies.
Ready-made camera UI
const scanner = ocr.openScanner(document.getElementById('scan-box'), { mode: 'magic' });
const { pages, pdf, cancelled } = await scanner.done; // pages: canvases, pdf: Blob
if (!cancelled) {
download(pdf, 'scan.pdf');
const text = (await ocr.scan(pages[0])).text; // and read it too
}The camera needs HTTPS (or localhost) and the page's Permissions-Policy must allow camera. It shows a live outline of the detected page, a Capture button, an Upload photo fallback, a style picker and a page strip.
Building blocks
const { canvas, found, corners } = await ocr.cropDocument(photoFile); // find the page, flatten it
const clean = await ocr.enhance(canvas, 'magic'); // 'magic' | 'bw' | 'gray' | 'color'
const pdf = await ocr.toPdf([clean, other], { title: 'Invoice' }); // Blob (application/pdf)
const read = await ocr.scan(photoFile, { autoCrop: true }); // crop first, then OCR| Call | What it does |
|---|---|
cropDocument(input, { corners? }) | Finds the page edges and corrects perspective. found is false when no page was detected (the whole image is returned). Pass your own corners to override. |
enhance(canvas, mode) | magic removes shadows and uneven light (best default), bw is high-contrast black and white, gray is stretched grayscale, color keeps colour but evens the lighting. |
toPdf(canvases, { quality, dpi, title }) | Builds a PDF with one image per page. The number of pages is limited by your plan (Free 5, Pro 50, Enterprise unlimited) and throws PLAN_LIMIT beyond it. |
openScanner(container, options) | The camera interface above. Options: mode, title, onDone. |
scan(input, { autoCrop: true }) | Runs cropDocument before recognition, useful for photos taken with a phone. |
For best results put the page on a surface that contrasts with it and keep all four corners in the frame. Tilted pages are fine; heavy folds, curled books and very low light are not handled.
Document type classifier
Know what was photographed before reading it. qovox-classifier.js (load it after the SDK) tells an Egyptian national ID, a passport and a driving licence apart, and says what is wrong with the picture when it cannot.
<script src="qovox-ocr-sdk.js"></script>
<script src="qovox-classifier.js"></script>
<script>
await QovoxClassifier.warmup(); // once at page load, so the first call is not the slow one
const r = await QovoxClassifier.classify(videoOrCanvasOrBlob, { ocr }); // ocr is optional, see below
// { doc_type: 'egypt_id' | 'passport' | 'drivers_license' | 'unknown', confidence: 0.94,
// bounding_box: [x, y, w, h], corners, quality: { blur, brightness, ok }, candidates, stages: { fast_ms, refine_ms },
// feedback: { code: 'BLURRY', message: '...', message_ar: '...' } }
</script>How it decides. Stage 1 (about 20 to 60 ms, pixels only) finds the page, flattens it, measures its shape and looks for the two dense text lines of a passport's machine-readable zone. That alone identifies passports. A card-shaped document is an ID or a licence, and telling those apart from pixels is unreliable, so stage 2 (only when you pass ocr, about 1 to 2 s) reads the text and looks for the words and a valid national number. Without ocr a card returns unknown with candidates and the code NEEDS_TEXT_CHECK.
Feedback codes (English and Arabic text included): BLURRY, TOO_DARK, GLARE, NO_DOCUMENT, TOO_SMALL, CUT_OFF, UNSUPPORTED, NEEDS_TEXT_CHECK, OK. Thresholds can be tuned with options.thresholds. Confidence is a heuristic score, not a calibrated probability.
Guided camera (manual framing)
Instead of guessing where the document is, the user picks its type and lines it up inside a fixed frame. Only the framed region is cropped, at the camera's full resolution, and checked on the device before OCR runs. Nothing is uploaded.
<script src="qovox-ocr-sdk.js"></script>
<script src="QovoxCameraOverlay.js"></script>
<script>
const ocr = await QovoxOCR.init({ licenseKey: 'QVX-...' });
const cam = QovoxCameraOverlay.open(document.getElementById('host'), { ocr, lang: 'auto' });
const res = await cam.done; // { canvas, ocr, type, validation, cancelled }
if (!res.cancelled) console.log(await ocr.scan(res.ocr));
// validation on its own (works on any canvas)
QovoxCameraOverlay.validateDocumentType(canvas, 'egypt_id', { margin: 0 });
// -> { ok, score, code: 'OK'|'ASPECT'|'EDGES'|'BLANK'|'FLAT_COLOR'|'LOW_TEXT'|'NO_FACE'|'NO_MRZ', checks, message, message_ar }
</script>Types: egypt_id, passport (with an MRZ guide), drivers_license, card. Live cues turn the frame green, orange or red for light, glare, blur and edge alignment, and the photo is taken by itself after about a second of steady green (or with the Capture button). Checks: shape, document edges against the frame, printed-text density, flat-colour/blank surface, a portrait where the photo should be (the browser's FaceDetector when it exists, otherwise a skin-tone heuristic) and the two MRZ lines on passports. They are heuristics that reject obvious mistakes; they do not prove a document is genuine. After two rejected tries the user may choose "Use anyway". res.ocr is grayscale with a contrast stretch; the 'bw' adaptive-threshold mode exists but lowered accuracy in our tests, so use it only for faint print.
Document engine V2 (IDs, passports, licences)
A field-level pipeline on top of the SDK: flatten the card, remove shadows, read each line separately, validate, and report a confidence per field.
<script src="qovox-ocr-sdk.js"></script>
<script src="QovoxEngineV2.js"></script>
<script>
const ocr = await QovoxOCR.init({ licenseKey: 'QVX-...' });
const r = await QovoxEngineV2.scanDocument(file, { docType: 'EGYPT_ID', ocr }); // 'EGYPT_ID' | 'PASSPORT' | 'DRIVERS_LICENSE'
r.fields.idNumber // { value, confidence, needsReview, validation, source }
r.review // field names under 85% confidence or failing validation: show these to the user
r.values // plain { field: value }
r.processed.canvas // the cleaned image that was read (use debugOverlay: true for a line-box overlay)
</script>What it does. Perspective-corrects the card (slightly widened so an edge in shadow is not cut), then builds image variants (lighting-corrected grayscale, Sauvola binarisation, CLAHE). It segments text rows, reads Arabic text rows with the Arabic model, and reads number rows with DigitNet, a small glyph classifier for Arabic-Indic digits that the general Arabic model reads poorly. Several reads of the national number are voted digit by digit. Arabic names are cleaned (ى/ي, ة/ه, hamza) against a list of common names. Passports use the ICAO check digits of the machine-readable zone.
Limits. The recogniser is still Tesseract in WebAssembly; V2 improves what goes in and checks what comes out. The Egyptian national number is validated by structure (century, real birth date, governorate, gender digit); there is no published check digit, so none is verified. Name and address row splitting is a heuristic tuned on synthetic cards. Dates are the weakest field. Always show the fields in review to a person.
Global capture flow (country + document)
One component for the whole journey: the user picks a country and a document type, the camera shows that document's guide (ID-1 cards, ID-3 passport with an MRZ box, A4 or Letter birth certificate), the frame turns green when steady and aligned, the photo is cropped at full camera resolution, and the country's rules read it.
<script src="qovox-ocr-sdk.js"></script>
<script src="QovoxCameraOverlay.js"></script><script src="QovoxOverlayManager.js"></script>
<script src="QovoxEngineV2.js"></script><script src="QovoxGlobalParser.js"></script><script src="QovoxCaptureFlow.js"></script>
<script>
const ocr = await QovoxOCR.init({ licenseKey: 'QVX-...' });
const flow = QovoxCaptureFlow.open(document.getElementById('host'), { ocr });
const { capture, parsed, country, docType } = await flow.done; // parsed.values, parsed.review, parsed.fields[x].confidence
// pieces on their own
QovoxOverlayManager.svg('AE', 'passport'); // guide drawing
await QovoxGlobalParser.parse(canvas, { country: 'SA', docType: 'national_id', ocr });
</script>Countries: Egypt, Saudi Arabia, UAE, USA, UK, Germany, France, Turkey, Russia, India and a Global option. Rules: Egyptian IDs (structure of the 14-digit number) and every passport (ICAO check digits) use QovoxEngineV2; Saudi (10 digits, starts with 1 or 2) and Emirates ID (15 digits, 784 prefix) use a Luhn-style check digit that public validators use but the issuers do not publish, so a mismatch flags the field instead of rejecting it. Other countries get the right language set, dates, a labelled name and the full text, not an ID-number rule. The guide shows real sizes; the drawn photo and number areas are approximate layouts.
Front and back scanning, merged record, output formats
ID cards and driving licences carry data on both sides. QovoxCaptureFlow now runs a state machine: front (portrait, name, national number) then back (address, job, marital status, issue and expiry dates, serial). The back camera opens by itself as soon as the front is captured, while the front is already being read. The two reads are merged into one record; a field printed on both sides must agree, otherwise it is listed in conflicts and flagged for review. Passports are one side.
<script src="QovoxOutputFormats.js"></script> <!-- plus the scripts from "Global capture" -->
<script>
const run = QovoxCaptureFlow.scanFrontAndBack(host, { ocr, country: 'EG', docType: 'national_id' });
const { record, format, output } = await run.done;
// record.values -> { fullName, idNumber, dateOfBirth, address, occupation, maritalStatus, issueDate, expiryDate, serialNumber }
// record.conflicts, record.review, record.photo (canvas), record.rawText.front / .back
// states: idle > front > back > processing > format > done (onState callback)
QovoxEngineV2.normalizeArabicNumerals('٢٩٨٠١٠١٠١٢٣٤٥٦') // '29801010123456'
</script>Digits. normalizeArabicNumerals() maps Arabic-Indic (٠-٩) and Persian (۰-۹) digits to 0-9 and the Arabic decimal and thousands marks to . and , before any regex or validator runs; every value the parser returns has been through it. Reading the digits correctly in the first place is a separate problem, handled by the glyph classifier in QovoxEngineV2. Formats. A dialog offers Fields JSON, Clean page (a card with the portrait; PNG download), Markdown and Plain text (the raw text of both sides). Copy and download are built in. Limits. The back-of-card reader is label driven and was tested on synthetic cards; address lines are the weakest part (street numbers and a second line can be missed), so expect them in the review list.
Number repair. QovoxGlobalParser.egypt.heal() repairs a misread national number: look-alike letters (O, I, |, l, V, S, B, Z), up to two look-alike digit swaps (0/5, 1/7, 7/8, 2/3, 4/6), one missing or one extra digit, keeping only candidates whose century digit, birth date and governorate code are valid. A repaired number is always flagged for review, and when several repairs are equally likely it is not applied and the options are returned as suggestions. The last digit is not verified: no check-digit rule is published. testChecksumHypothesis([...]) lets you test one candidate rule on real numbers locally. Long digit strings under the barcode go to barcodeText and never become the serial number or the national number.
API reference
QovoxOCR.init(options) → Promise<QovoxOCR>
| Option | Default | Meaning |
|---|---|---|
licenseKey | required | Your key. QVX-DEMO-… keys skip the check (used by the website demo). |
languages | 'auto' | 'auto' starts with one light pass (ara+eng, which also reads Latin text) and only downloads heavier models (Russian, Chinese, Japanese, Korean, Hindi, Hebrew, Greek, Thai, European pack) when that result looks wrong. It remembers what worked for the next scan. Or give any Tesseract code(s): 'ara', 'chi_sim+eng'. All 100+ are in QovoxOCR.languages. Naming the language is always faster and more accurate than 'auto'. |
autoTries | 5 | How many language sets 'auto' may try before keeping the best. Raise it for rarer scripts. |
maxSide | 2000 | Longest image side in px. Larger images are shrunk (aspect ratio kept) before recognition, which cuts time without hurting accuracy. Small images are enlarged to about 1500 px. |
cache | true | Keeps downloaded language models in the browser's IndexedDB so each model is downloaded once. Set false to disable. |
maxPages | 10 | Pages read from a PDF. |
telemetry | true | Anonymous page counter used for your dashboard and quota. No document data. |
apiEndpoint | Qovox API | Override for tests or a private deployment. |
workerPath, corePath, langPath | CDN | Self-hosted engine files, see below. |
ocr.scan(input, options?) → Promise<Result>
input: File, Blob, canvas, <img>, ImageBitmap, ArrayBuffer or a same-origin URL. Options: documentType ('auto' | 'invoice' | 'id' | 'passport'), languages, onProgress({status, progress}). Scans on one instance run one at a time. Image preparation (resize, grayscale, contrast) and text recognition each run in their own Web Worker, so the page stays responsive.
ocr.preload(languages?)
Optional. Downloads and warms the engine before the first scan (for example when the user opens a file picker). Resolves true when ready.
ocr.destroy()
Stops the engine and frees its memory. Call it when your component unmounts.
Errors
Errors are QovoxError with a code (PLAN_LIMIT when a PDF has more pages than the plan allows): NO_LICENSE_KEY, KEY_FORMAT, KEY_UNKNOWN, KEY_REVOKED, EMAIL_NOT_VERIFIED, DOMAIN_NOT_ALLOWED, QUOTA_EXCEEDED, UNSUPPORTED_FILE, ENGINE_LOAD_FAILED.
REST API
You do not need to call these yourself, the SDK does, but they are stable and useful for server-side checks and dashboards. Base URL: https://qovox-ocr-api.qovox222.workers.dev/api/v1. Responses are JSON and CORS is open (Access-Control-Allow-Origin: *).
Public endpoints (no login)
| Endpoint | Purpose |
|---|---|
GET /health | { "ok": true, "status": "online" } |
POST /verify-license | Body { "licenseKey", "domain" }. The browser's Origin header wins over domain, so a page cannot claim another site. |
POST /heartbeat | Body { "licenseKey", "domain", "scans": 3 }. Adds pages to this month's counter (max 1000 per call). |
// POST /verify-license -> 200
{ "valid": true, "plan": "pro", "scansUsed": 1280, "scansLimit": 25000, "pdfPages": 50 }
// not allowed -> 200 (the HTTP status is 200, check "valid")
{ "valid": false, "code": "DOMAIN_NOT_ALLOWED", "reason": "shop.example.com is not allowed for this key. Add it in your Qovox OCR dashboard." }
code | Meaning |
|---|---|
KEY_FORMAT | Not a Qovox key (QVX-XXXX-XXXX-XXXX-XXXX). |
KEY_UNKNOWN | No such key (or it was rotated). |
KEY_REVOKED | Revoked in the dashboard. |
DOMAIN_NOT_ALLOWED | The page's domain is not on the key's list. localhost always passes. |
QUOTA_EXCEEDED | The month's pages are used up. scansUsed and scansLimit are included. |
SUBSCRIPTION_ENDED | The paid subscription ended (or was cancelled). Only what the Free plan includes keeps working: the oldest key, one domain per key, 500 pages. Renew to restore the rest. |
PLAN_LIMIT | The key or domain is beyond what the current plan includes. |
EMAIL_NOT_VERIFIED | The key owner has not verified their email. |
ACCOUNT_FROZEN | The account is suspended. |
Account endpoints (used by the dashboard)
These need Authorization: Bearer <Firebase ID token> from a signed-in user with a verified email; otherwise they answer 401 UNAUTHENTICATED or 403 EMAIL_NOT_VERIFIED.
| Endpoint | Purpose |
|---|---|
GET /me | Account, plan, limits, keys and usage history. |
POST /licenses | Create a key ({ label, domains }). The key is returned once and only a hash is stored. |
PATCH /licenses/:id | Change domains, label or status. |
POST /licenses/:id/rotate | Issue a new key; the old one stops working. |
DELETE /licenses/:id | Delete a key and its usage. |
POST /upgrade-request | Ask for Pro or Enterprise. |
Subscriptions
Paid plans are monthly subscriptions. The plan is live while the subscription period is paid. If a renewal payment fails the plan stays on for a 3-day grace period; after that, or as soon as you cancel, the account falls back to the Free limits automatically: the oldest key and one domain per key keep working, everything else answers SUBSCRIPTION_ENDED, and localhost always works. Renewing restores everything within seconds. The license response carries subscriptionEnded: true while an account is in this state, so you can warn your own admins.
Licensing and domains
Each key lists the domains where it may run (app.example.com or *.example.com on paid plans). localhost always works. init() sends the key and the page origin to POST /api/v1/verify-license; the server answers { valid, plan, scansUsed, scansLimit }. Keys can be rotated or revoked in the dashboard. Treat a key like a public identifier: it is domain-locked, so copying it to another site does not work.
Self-hosting the engine
By default the engine (Tesseract.js, WebAssembly) and language data load from a public CDN and are cached. To serve them yourself, copy worker.min.js, the tesseract-core*.wasm.js files and your *.traineddata.gz files to your site and pass workerPath, corePath and langPath. Also set tesseractUrl to your copy of tesseract.min.js, and pdfjsUrl / pdfjsWorkerUrl if you scan PDFs.
Privacy
Images and PDFs are decoded and recognised in the visitor's browser. Network calls to Qovox contain the license key, the page origin, the SDK version and a page count. Nothing from the document is transmitted.