AI Projects · Technical writing for developers
Reading OCR Confidence
A short guide for developers deciding which OCR results can go straight into a system and which need a person to check them, built from the example output in Tesseract's own documentation.
It includes an evidence-to-decision table and a small JavaScript review rule, tested against documented and synthetic rows. Researched and drafted with AI assistance; How I Build With AI explains the process.
When OCR Needs a Second Look
Turning Tesseract's TSV word confidence into review decisions, without treating a high score as an approval
Independent writing sample. It works from example output published in the Tesseract documentation; no OCR was run and no engine was benchmarked for it.
Two words from one documented image show the problem.
In the TSV example on Tesseract's command-line usage page, the word $43,456.78 comes back with a confidence of 93.263031. In the hOCR example for the same image, the word preguigoso. comes back with x_wconf 92 [1].
The amount is right: the image reads $43,456.78. The second word is wrong. The image reads preguiçoso, Portuguese for "lazy," and the documented command ran with the English model only (-l eng) [1].
These are documentation examples, not a benchmark. A rule that accepts every word scoring 90 or more takes the wrong word along with the right one. A rule that demands 95 sends both to a person, including the correct amount. Neither rule knows which value is true. That is the job of the review step around the OCR, and it is what this piece is about.
What the TSV gives you
Add tsv to the end of a Tesseract command and you get one row per recognized element, with twelve columns: level, page_num, block_num, par_num, line_num, word_num, left, top, width, height, conf and text [1][2]. The manual page sums up the format in one line: a row for each recognized word, carrying its bounding box, confidence and text [3].
Not every row is a word. In the documented sample, rows at levels 1 to 4 line up with the page, block, paragraph and line numbering, carry a conf of -1 and have no text. Only level 5 rows carry a word and a score [1]. Treat every row as a candidate value and the first one you meet is the whole page, with no text and a score of -1.
The score also looks different by format: the documented TSV example prints six decimal places, the hOCR example whole numbers. Both ran on the same image with -l eng, and the hOCR header names a Tesseract 5.0.1 build [1]. Pick one output as your source of truth before comparing anything with a threshold.
What the score doesn't tell you
Neither the usage page nor the manual page describes conf as the probability that a word is correct [1][3]. Treat it as a signal for where to look first, not as a measurement of truth.
Two more gaps matter in production. First, a word is not a field. $43,456.78 happens to be one word, but a date such as 12 Oct 2026 arrives as three word rows. If one of them is misread, or the code that joins them picks up a word from the next column, the assembled date is wrong even when every score looks high. Check the assembled field, not only each word. Second, any threshold is your policy, not Tesseract's advice. The 95 used below is an illustration chosen for this example.
From evidence to decision
The table separates what a TSV row shows from what a hypothetical review policy does with it.
| What the row shows | What the review rule does | Why |
|---|---|---|
Level 1 to 4, conf of -1, no text | Skip it as a value; keep it for layout | It describes structure, not a word [1] |
| Level missing or not 1 to 5 | Send to review | The rule doesn't guess what an unknown row is |
| Level 5, score of 95 or more, field check passes | Propose the value with its raw text, position and the policy that passed it | The policy is met, but acceptance still belongs to a person or a later step |
| Level 5, score below 95 | Send to review | The correct $43,456.78 at 93.26 lands here |
Level 5, score missing, not a number or -1 | Send to review | Missing metadata is not a pass |
| Any score, field check fails | Send to review | A high score can't overrule a failed check: a synthetic $1,25O.00, with a letter O, at 98.1 fails the amount check |
| A script or model suggests a corrected value | Show it beside the raw text for approval | A normalization is a proposal, not a silent fix |
A small rule you can test
The function below applies that table to one parsed TSV row, in plain JavaScript with no dependencies. The value stays text, so nothing is rounded. Every result carries source (page, block, paragraph, line, word and box, with null where the row has none) and the policy that made the call, so any proposal or review item traces back to its row. checkField is your field's validation. This one checks a US-style currency amount and knows nothing about spelling.
// Review rule for one parsed Tesseract TSV row (all fields as strings).
// POLICY is this example's choice, not a Tesseract recommendation.
const POLICY = { id: 'ocr-review-example-v2', minConf: 95 };
// Where the row came from. A field the row doesn't have stays null.
function sourceOf(row) {
const get = (k) => {
const v = row[k] == null ? '' : String(row[k]).trim();
return v === '' ? null : v;
};
const box = ['left', 'top', 'width', 'height'].map(get);
return {
page: get('page_num'), block: get('block_num'), par: get('par_num'),
line: get('line_num'), word: get('word_num'),
box: box.includes(null) ? null : box,
};
}
function reviewRow(row, checkField) {
const base = { policy: POLICY.id, minConf: POLICY.minConf, source: sourceOf(row) };
if (['1', '2', '3', '4'].includes(row.level)) {
return { ...base, action: 'skip', reason: 'structural row' };
}
const raw = row.text ?? '';
if (row.level !== '5') return { ...base, action: 'review', reason: 'unknown level', raw };
if (raw.trim() === '') return { ...base, action: 'review', reason: 'no text', raw };
const confText = row.conf == null ? '' : String(row.conf).trim();
const conf = Number(confText);
if (confText === '' || !Number.isFinite(conf) || conf < 0) {
return { ...base, action: 'review', reason: 'confidence missing', raw };
}
const check = checkField(raw);
if (!check.ok) return { ...base, action: 'review', reason: check.reason, raw, conf };
if (conf < POLICY.minConf) {
return { ...base, action: 'review', reason: 'below policy threshold', raw, conf };
}
return { ...base, action: 'propose', raw, value: check.value, conf };
}
function checkAmount(raw) {
if (!/^\$(\d{1,3}(,\d{3})+|\d+)\.\d{2}$/.test(raw)) {
return { ok: false, reason: 'not a currency amount' };
}
return { ok: true, value: raw.slice(1).replace(/,/g, '') };
}
The function was tested in a browser JavaScript engine against three rows from the documentation and eight synthetic rows. Every result was also checked against its input: same raw text, same position, and the policy named. That tests the review logic, not OCR.
- Documented level 1 row,
confof-1: skip, structural row. - Documented
$43,456.78at93.263031: review, below policy threshold. - Documented
(quick)at95.965691, run through the amount check: review, not a currency amount. - Synthetic
$1,250.00at97.5: propose, value1250.00. - Synthetic
$1,25O.00(a letter O) at98.1: review, not a currency amount. - Synthetic word rows with an empty score or a score of
-1: review, confidence missing. - Synthetic word row with a score of
96and no text: review, no text. - Synthetic rows with level
6and with no level: review, unknown level. - Synthetic
$75.00at97with no box: propose, value75.00, boxnull.
Before a value enters your system
- Keep the raw text, the score and the position with every proposed value, so a reviewer can find the word on the page. If the row has no box, record that it's missing.
- Record which rule and threshold made the call, so a later change can be traced.
- Treat a missing or
-1score on a word row as a reason to look. - Validate the field, not only the word.
- Let a model suggest a normalized value only beside the original, and let a person approve the change.
- Decide separately what "accepted" means downstream. This rule only proposes.
Confidence is still useful. It tells you where to look first. It can't tell you what the document says. The image, the field rules and a reviewer settle that, and a good pipeline makes the uncertain part visible instead of quietly passing it through.
References
- Tesseract documentation, "Command Line Usage": TSV and hOCR output examples and the example image eurotext.png. tesseract-ocr.github.io/tessdoc/Command-Line-Usage.html. Checked October 5, 2026. ↩
- The same page's Markdown source on GitHub. github.com/tesseract-ocr/tessdoc/blob/main/Command-Line-Usage.md. ↩
- Tesseract manual page, tessedit_create_tsv entry. github.com/tesseract-ocr/tesseract/blob/main/doc/tesseract.1.asc. ↩