Two Codex runs, configured with different model identifiers, agreed on the wrong price. Both returned structured output without flagging uncertainty.
I was testing extraction for Wombat because I wanted capital-gain and loss calculations I could trust. Changing the OCR engine could recover more numbers. I still needed to establish whether those numbers belonged to the right transactions.
I tested 78 distinct financial PDFs across 18 institution groups, covering all 531 physical pages. Fifteen groups had five originals each. Three had only one available original each; I didn’t count copies as extra samples. The corpus mixed statements, confirmations, cards, loans and retirement documents. These were mostly digital PDFs, sometimes with image-only regions. This says little about damaged scans or handwriting.
Benchmark numbers
Measured on an Apple M3 Mac with 16 GiB of RAM on October 9, 2026. One traversal per configuration; desktop activity and system swap weren’t isolated. These are observed workflow latencies, not repeated-run performance estimates. The scores use a source-derived reference corrected after examining disagreements; they do not measure verified capital-gain accuracy.
What the percentages mean
The unit is a decimal occurrence on a physical page. A repeated value counts again each time it appears. A match preserves the digits, printed decimal precision and the benchmark’s sign convention. These percentages do not measure dollars recovered or the correctness of a gain calculation.
| Metric | Calculation | What it tells me |
|---|---|---|
| Recall | Matched occurrences ÷ expected reference occurrences × 100 | How much of the reference did the extractor recover? Missing values lower recall. |
| Precision | Matched occurrences ÷ emitted occurrences × 100 | How much of the extractor’s output matched the reference? Extra or wrong values lower precision. |
| Exact PDF rate | PDFs with exact inventories and passing protocols ÷ attempted PDFs × 100 | How often did a whole PDF match on every page? One missing, extra or changed occurrence makes that PDF non-exact. |
For an invented example, suppose the reference has 100 occurrences. The extractor emits 95, of which 90 match. Recall is 90 ÷ 100 = 90%. Precision is 90 ÷ 95 = 94.74%. Ten expected occurrences were not matched, and five emitted occurrences were not matched. A changed digit can hurt both scores: the expected value is missing and the wrong value is extra.
Matches are counted within each physical page, then pooled across the workflow. A document containing many decimal occurrences contributes more to recall and precision than a short document. The exact PDF rate gives each attempted PDF one vote. Even a perfect inventory can put basis on the wrong transaction row; this benchmark does not check row ownership.
Codex recovered 99.942% of expected decimal occurrences in 43.17 minutes, with exact inventories on 67 of 78 PDFs. Direct Tesseract PSM 3 took 14.89 minutes and had 38 exact PDFs. Embedded text took 0.10 minutes, but it supplied most of the reference and is not an independent accuracy result.
The time bars include unsuccessful attempts. Rendering is included; Docling model warmup is excluded. These are single observed traversals on the same corpus.
An exact inventory requires the same decimal multiset on every physical page and a passing protocol. It does not check that the numbers belong to the right financial rows.
The scatter plot uses the exact PDF rate on its vertical axis. Codex’s point is 67 ÷ 78 × 100 = 85.90%, not its 99.942% decimal recall. Higher means more exact PDFs; farther left means less elapsed time. The table keeps decimal recall and precision separate from whole-PDF exactness.
| Workflow | Raw decimal recall | Raw decimal precision | Protocol passes | Exact PDFs | Attempted minutes |
|---|---|---|---|---|---|
| Embedded PDF text* | 99.903% | 100.000% | 78/78 | 73/78 | 0.10 |
| Docling, native text | 99.749% | 99.807% | 78/78 | 65/78 | 10.72 |
| Tesseract, PSM 3 | 98.388% | 99.677% | 78/78 | 38/78 | 14.89 |
| Tesseract, PSM 6 | 96.775% | 99.514% | 78/78 | 33/78 | 13.64 |
| OCRmyPDF, force OCR | 85.596% | 99.473% | 76/78 | 39/78 | 10.54 |
| Docling, full-page Tesseract | 99.566% | 99.806% | 78/78 | 49/78 | 18.80 |
| Docling, full-page RapidOCR | 99.759% | 99.518% | 78/78 | 56/78 | 38.83 |
| Codex, gpt-6.1-sol / medium | 99.942% | 99.865% | 77/78 | 67/78 | 43.17 |
Embedded text isn’t OCR. Most of that row’s agreement is circular because it produced the text reference. The added source-image values expose known omissions in that layer.
For Codex, 10,352 of 10,358 expected occurrences matched, giving 99.942% recall. It emitted 10,366 decimal occurrences, giving 10,352 ÷ 10,366 = 99.865% precision. That leaves six unmatched expected occurrences and fourteen unmatched emitted occurrences under this checker. These are occurrence counts, not a count of bad transactions or incorrect tax calculations.
Protocol passes count PDFs whose pipeline completed and whose output met its required format. A protocol pass does not establish numerical or financial correctness. An exact PDF needs both a protocol pass and an exact inventory.
Raw recall and precision include valid decimal output from unsuccessful attempts. A failed pipeline or malformed response remains a failure; those values aren’t an accepted ingestion result. The Codex run completed its calls, but one PDF’s response included an integer in the required decimal-only array. JSON shape validation alone didn’t catch that contract error.
Tested versions and settings
These are the versions installed for the October 9, 2026 benchmark, not claims about the latest releases.
- Machine: Apple M3, 16 GiB RAM; local CPU workflows.
- Poppler 26.09.0:
pdftotext -layout; embedded-text baseline, not OCR. - Tesseract 5.5.3: English, 300 dpi grayscale, PSM 3 and PSM 6 as separate runs;
OMP_THREAD_LIMIT=2. - OCRmyPDF 17.13.0: force OCR, two jobs, PSM 3, optimization disabled; wraps Tesseract.
- Docling 2.135.0: native-text, full-page Tesseract and full-page RapidOCR as three separate workflows; tables enabled, remote services disabled, roughly 216 dpi for full-page OCR.
- RapidOCR 3.9.1: local PP-OCRv6 detection and recognition, plus the
ch_ppocr_mobile_v2.0orientation classifier. - ONNX Runtime 1.31.0: runtime for the RapidOCR models.
- Codex CLI 0.162.0: existing ChatGPT login; 180 dpi color pages, at most four pages per chunk; 152 sequential calls for the full corpus.
- Requested full-corpus model:
gpt-6.1-sol, medium reasoning. - Requested comparison model:
gpt-6-astra, medium reasoning, on eight originals only. The clean Sol comparison usedgpt-6.1-sol/ medium, with project instructions, memory, plugins and app tools disabled per invocation.
The Codex CLI did not expose a resolved backend model snapshot. The identifiers above describe what I requested.
What I measured
Every workflow attempted every selected original in full. I asked Codex to transcribe every visible decimal literal, preserving precision and repeated occurrences. It received page images, not the reference text. The local workflows produced text or table cells that I normalized into the same decimal inventory.
The checker compared a multiset per physical page. Missing values, extra values, changed digits and changed precision affected the score. Integers, dates without decimal points, ordinary text, reading order and ownership of values were outside the metric.
The reference came from the PDF’s embedded text, with ten values added after checking source images. It contained 10,358 decimal occurrences, including 1,349 zeros. Recall counts recovered expected occurrences; precision counts how many emitted occurrences match the reference. It was not an independently labeled human gold set. I corrected the derived checker after examining disagreements, so these are exploratory agreement figures. I haven’t manually audited all 531 pages. The PDFs and raw evidence remain private.
A parenthesis rule was also a convention: the prompt encoded parenthetic decimals as negatives, including explanatory disclosures where the financial meaning wasn’t negative. The private report separates magnitude-only recovery from that signed convention.
Codex timings include page rendering, CLI startup, network time, inference and output serialization. I recorded token usage, but this wasn’t a per-token API billing experiment. The event audit recorded no tool calls during transcription.
PSM 3 automatically segments a page; PSM 6 assumes a single block. Those layout assumptions are described in Tesseract’s quality guide. I charged rendering to each independently usable configuration.
OCRmyPDF wraps Tesseract and PDF processing. Docling adds layout and table reconstruction; its full-page OCR configuration is another workflow, not another independent recognizer. Its Tesseract path also performed orientation detection. Docling model warmup was logged separately. Two threads per configured model or stage did not mean a global two-core cap: pipelined stages could overlap.
Raster resolutions differed across workflows. I didn’t test every OCR product, GPU acceleration or arbitrary Codex model configurations.
I also kept timings separate by institution. A short card statement and a dense investment statement create different work; an overall average conceals that mix. The institution tables below include document counts, page counts, median whole-PDF latency, numeric agreement and failures. The selection favored available recent originals across document families; some templates were older. This was a purposive sample, not a random sample of financial PDFs.
The checker needed debugging too
Some apparent model errors were reference errors. A chart legend existed only as an image. Payment amounts were spread across character boxes, so the text layer didn’t contain a contiguous decimal. Dot leaders obscured printed minus signs. A date separator looked like a negative quantity to my regex.
I preserved the original text and candidate outputs, versioned the checker corrections, and applied each correction to every workflow. Those changes don’t make the benchmark blind or independent. They make the remaining comparison more useful.
A PDF can contain plenty of text and still omit visible numbers from its text layer. I can’t use “text exists” as proof that native extraction is complete.
The wrong numbers were plausible
The primary Codex run changed one purchase-price digit and misread a cash amount in two repeated fields. The repeated fields agreed with each other, so an equality check would have passed. Quantity multiplied by price caught the cash discrepancy, allowing for printed precision and cash rounding. That check concerns printed trade cash, not tax-adjusted basis; fees and source rounding need their own handling.
A separate eight-original comparison requested gpt-6-astra at medium reasoning. Both configurations had exact numeric inventories on the preselected five short originals. On three additional stress originals, the second configuration corrected the cash reading but repeated the wrong price. Those stress cases were selected after seeing a failure, so they aren’t unseen validation data. I can’t call model agreement an independent truth check or infer a global best model from this subset.
All three configurations used medium reasoning. The clean Sol configuration disabled project instructions, memory, plugins and app tools for each invocation. That changed the configuration, not just the requested model. These single traversals do not establish a causal speed improvement.
| Subset | Configuration | Exact PDFs | Signed matches | Total seconds |
|---|---|---|---|---|
| Five short originals | Codex Sol | 5/5 | 160/160 | 70.67 |
| Five short originals | Codex Astra | 5/5 | 160/160 | 99.80 |
| Five short originals | Codex Sol, clean config | 5/5 | 160/160 | 69.59 |
| Three failure-selected originals | Codex Sol | 1/3 | 2448/2451 | 286.56 |
| Three failure-selected originals | Codex Astra | 2/3 | 2450/2451 | 239.38 |
| Three failure-selected originals | Codex Sol, clean config | 2/3 | 2449/2451 | 243.22 |
The short subset covered SoFi, Wealthfront, Robinhood, Chase and Citi: 15 pages and 160 expected decimal occurrences. The additional stress subset covered SoFi, Wealthfront and Acorns: 31 pages and 2,451 occurrences. Every response in these subsets passed the protocol. The larger Astra comparison was not run across all 78 originals.
I retried the troublesome page three times at 180 dpi and three times at 300 dpi. The 180 dpi page passed twice; the 300 dpi page passed all three times. A 600 dpi crop of the source row also passed three times. This is evidence for a targeted higher-resolution retry, with a tiny sample and a changed single-page context. It isn’t a guarantee.
| Retry input | Exact inventories |
|---|---|
| Full page, 180 dpi | 2/3 |
| Full page, 300 dpi | 3/3 |
| Source row crop, 600 dpi | 3/3 |
The primary run reported no uncertainty entries anywhere, despite the confirmed mistakes. Confidence wasn’t a useful acceptance test here.
A perfect inventory can still produce wrong gains
This example uses invented transactions, not my account data:
| Extraction | Row | Given category | Proceeds | Basis | Gain |
|---|---|---|---|---|---|
| Correct | Alpha | Long term | 120 | 100 | 20 |
| Correct | Beta | Short term | 80 | 60 | 20 |
| Basis rows swapped | Alpha | Long term | 120 | 60 | 60 |
| Basis rows swapped | Beta | Short term | 80 | 100 | -20 |
Both versions contain the same numbers, and both total 40 in gains. The category results differ. I ran this invented row swap through the actual inventory-matching function; it passed all four numeric occurrences. This demonstrates a limit of the benchmark, not a measured error from an OCR engine or a Wombat financial-parser test.
Tests for gains need to assert row ownership, account-specific lots, source basis, categories and final numbers. A numeric inventory or balanced total can’t replace those assertions.
Rendering can fail before recognition
The fixed OCRmyPDF workflow failed on two originals. One renderer couldn’t create a bitmap; another run produced an invalid output PDF. Both original PDFs passed a syntax check, which doesn’t establish financial correctness.
Switching to Ghostscript didn’t resolve either failure. One retry attempted a 1.5-billion-pixel image and hit its image-size guard. I kept the guard enabled.
A separate recovery pipeline rendered every original page at a bounded 300 dpi, rebuilt a derivative image PDF, then ran OCRmyPDF. Both recovered attempts produced valid PDFs with all physical pages retained. They still missed some decimal values. These were failure-led retries on two sources, not replacement scores for the main corpus.
The recovered 73-page brokerage PDF took 75.37 seconds and matched 1,130 of 1,142 expected decimals. The five-page bank PDF took 9.43 seconds and matched 23 of 26. Those times include the compound recovery workflow. Successful rendering still needed an extraction check.
Results by institution
Median seconds are per successfully completed whole PDF, including rendering; Docling model warmup is excluded. Each workflow attempted every selected PDF. Recall includes expected values from failed attempts. Precision includes any decimal output those attempts emitted; that partial output remains diagnostic. Exact means a matching decimal multiset on every physical page, with a passing protocol. It does not mean every financial field is correct. Percentages are rounded to three decimals.
Document types and lengths differ. These medians cannot rank how difficult a bank is to parse, and a median can hide one long document: Robinhood included a 73-page original. Apple / Goldman Sachs includes Marcus-branded documents. Rocket Mortgage is grouped by document branding. American Express evidence was available as a spreadsheet, without a dedicated PDF sample here.
Acorns
1 original PDF, 23 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.21 | 100.000% | 100.000% | 1/1 | 1/1 |
| Tesseract PSM 3 | 33.78 | 100.000% | 100.000% | 1/1 | 1/1 |
| Tesseract PSM 6 | 31.28 | 100.000% | 100.000% | 1/1 | 1/1 |
| OCRmyPDF | 23.02 | 92.503% | 99.953% | 1/1 | 0/1 |
| Docling native | 56.60 | 100.000% | 100.000% | 1/1 | 1/1 |
| Docling Tesseract | 71.61 | 100.000% | 100.000% | 1/1 | 1/1 |
| Docling RapidOCR | 136.47 | 100.000% | 98.702% | 1/1 | 0/1 |
| Codex Sol | 250.12 | 99.956% | 99.956% | 1/1 | 0/1 |
Apple / Goldman Sachs
5 original PDFs, 17 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 5.97 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 6 | 5.49 | 100.000% | 100.000% | 5/5 | 5/5 |
| OCRmyPDF | 9.03 | 99.351% | 99.351% | 5/5 | 4/5 |
| Docling native | 2.05 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 6.71 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling RapidOCR | 13.72 | 100.000% | 99.355% | 5/5 | 4/5 |
| Codex Sol | 14.08 | 100.000% | 100.000% | 5/5 | 5/5 |
Capital One
5 original PDFs, 18 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.05 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 8.42 | 95.270% | 99.296% | 5/5 | 2/5 |
| Tesseract PSM 6 | 8.10 | 91.216% | 100.000% | 5/5 | 1/5 |
| OCRmyPDF | 5.73 | 91.892% | 100.000% | 5/5 | 0/5 |
| Docling native | 3.47 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 9.19 | 98.649% | 100.000% | 5/5 | 3/5 |
| Docling RapidOCR | 17.13 | 100.000% | 100.000% | 5/5 | 5/5 |
| Codex Sol | 15.42 | 100.000% | 98.013% | 5/5 | 2/5 |
Charles Schwab
5 original PDFs, 26 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.09 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 10.24 | 98.415% | 98.763% | 5/5 | 2/5 |
| Tesseract PSM 6 | 9.61 | 98.592% | 98.765% | 5/5 | 3/5 |
| OCRmyPDF | 10.62 | 98.592% | 98.765% | 5/5 | 2/5 |
| Docling native | 4.37 | 100.000% | 99.475% | 5/5 | 4/5 |
| Docling Tesseract | 12.75 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling RapidOCR | 30.06 | 99.648% | 100.000% | 5/5 | 3/5 |
| Codex Sol | 23.44 | 99.472% | 98.949% | 5/5 | 2/5 |
Chase
5 original PDFs, 20 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 4.37 | 99.401% | 100.000% | 5/5 | 4/5 |
| Tesseract PSM 6 | 4.05 | 98.802% | 98.214% | 5/5 | 3/5 |
| OCRmyPDF | 4.24 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling native | 2.99 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 6.12 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling RapidOCR | 12.23 | 100.000% | 98.817% | 5/5 | 3/5 |
| Codex Sol | 14.33 | 100.000% | 100.000% | 5/5 | 5/5 |
Citi
5 original PDFs, 26 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 9.22 | 98.413% | 100.000% | 5/5 | 3/5 |
| Tesseract PSM 6 | 8.55 | 96.032% | 100.000% | 5/5 | 2/5 |
| OCRmyPDF | 6.52 | 98.413% | 100.000% | 5/5 | 3/5 |
| Docling native | 4.03 | 99.206% | 100.000% | 5/5 | 4/5 |
| Docling Tesseract | 11.13 | 96.032% | 100.000% | 5/5 | 2/5 |
| Docling RapidOCR | 21.95 | 99.206% | 100.000% | 5/5 | 4/5 |
| Codex Sol | 21.34 | 100.000% | 100.000% | 5/5 | 5/5 |
Comenity
5 original PDFs, 24 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 9.36 | 95.337% | 98.925% | 5/5 | 0/5 |
| Tesseract PSM 6 | 8.65 | 90.674% | 96.685% | 5/5 | 0/5 |
| OCRmyPDF | 7.18 | 95.855% | 97.884% | 5/5 | 0/5 |
| Docling native | 3.11 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 9.03 | 98.964% | 99.479% | 5/5 | 4/5 |
| Docling RapidOCR | 17.87 | 98.964% | 100.000% | 5/5 | 3/5 |
| Codex Sol | 21.11 | 100.000% | 100.000% | 5/5 | 5/5 |
Discover
5 original PDFs, 18 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 4.69 | 87.023% | 96.610% | 5/5 | 0/5 |
| Tesseract PSM 6 | 4.52 | 84.733% | 93.277% | 5/5 | 0/5 |
| OCRmyPDF | 4.11 | 96.947% | 98.450% | 5/5 | 3/5 |
| Docling native | 3.28 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 7.89 | 97.710% | 98.462% | 5/5 | 3/5 |
| Docling RapidOCR | 15.58 | 99.237% | 100.000% | 5/5 | 4/5 |
| Codex Sol | 15.98 | 100.000% | 98.496% | 5/5 | 3/5 |
Fidelity
1 original PDF, 4 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.04 | 96.685% | 100.000% | 1/1 | 0/1 |
| Tesseract PSM 3 | 7.28 | 99.448% | 100.000% | 1/1 | 0/1 |
| Tesseract PSM 6 | 6.75 | 100.000% | 100.000% | 1/1 | 1/1 |
| OCRmyPDF | 5.45 | 96.133% | 98.305% | 1/1 | 0/1 |
| Docling native | 6.14 | 96.685% | 100.000% | 1/1 | 0/1 |
| Docling Tesseract | 11.05 | 100.000% | 100.000% | 1/1 | 1/1 |
| Docling RapidOCR | 24.69 | 99.448% | 100.000% | 1/1 | 0/1 |
| Codex Sol | 23.71 | 100.000% | 100.000% | 1/1 | 1/1 |
J.P. Morgan
5 original PDFs, 80 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.13 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 39.81 | 99.180% | 99.740% | 5/5 | 2/5 |
| Tesseract PSM 6 | 36.77 | 99.612% | 99.870% | 5/5 | 2/5 |
| OCRmyPDF | 26.11 | 97.454% | 99.122% | 5/5 | 2/5 |
| Docling native | 38.52 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 51.15 | 99.741% | 99.957% | 5/5 | 2/5 |
| Docling RapidOCR | 103.16 | 100.000% | 100.000% | 5/5 | 5/5 |
| Codex Sol | 63.17 | 100.000% | 100.000% | 4/5 | 4/5 |
Morgan Stanley
5 original PDFs, 26 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 5.66 | 99.627% | 99.627% | 5/5 | 3/5 |
| Tesseract PSM 6 | 5.42 | 93.657% | 99.735% | 5/5 | 3/5 |
| OCRmyPDF | 5.37 | 99.627% | 99.627% | 5/5 | 3/5 |
| Docling native | 1.80 | 99.627% | 100.000% | 5/5 | 4/5 |
| Docling Tesseract | 8.12 | 99.254% | 100.000% | 5/5 | 1/5 |
| Docling RapidOCR | 14.47 | 99.502% | 100.000% | 5/5 | 3/5 |
| Codex Sol | 18.83 | 100.000% | 100.000% | 5/5 | 5/5 |
Robinhood
5 original PDFs, 87 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 8.58 | 97.617% | 99.620% | 5/5 | 4/5 |
| Tesseract PSM 6 | 8.02 | 93.150% | 99.365% | 5/5 | 0/5 |
| OCRmyPDF | 6.27 | 14.520% | 100.000% | 4/5 | 2/5 |
| Docling native | 5.11 | 99.255% | 99.701% | 5/5 | 2/5 |
| Docling Tesseract | 13.52 | 99.181% | 99.701% | 5/5 | 1/5 |
| Docling RapidOCR | 22.83 | 99.255% | 99.701% | 5/5 | 2/5 |
| Codex Sol | 18.00 | 100.000% | 100.000% | 5/5 | 5/5 |
Rocket Mortgage
5 original PDFs, 10 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.04 | 97.647% | 100.000% | 5/5 | 1/5 |
| Tesseract PSM 3 | 4.29 | 90.588% | 98.718% | 5/5 | 1/5 |
| Tesseract PSM 6 | 4.12 | 89.412% | 97.436% | 5/5 | 1/5 |
| OCRmyPDF | 3.91 | 91.765% | 100.000% | 5/5 | 1/5 |
| Docling native | 3.19 | 97.647% | 93.785% | 5/5 | 1/5 |
| Docling Tesseract | 5.92 | 96.471% | 94.798% | 5/5 | 1/5 |
| Docling RapidOCR | 8.57 | 97.647% | 93.785% | 5/5 | 1/5 |
| Codex Sol | 17.96 | 100.000% | 100.000% | 5/5 | 5/5 |
ScholarShare 529
5 original PDFs, 18 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.04 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 4.90 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 6 | 4.57 | 100.000% | 100.000% | 5/5 | 5/5 |
| OCRmyPDF | 3.66 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling native | 2.65 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 5.93 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling RapidOCR | 9.75 | 100.000% | 100.000% | 5/5 | 5/5 |
| Codex Sol | 12.29 | 100.000% | 100.000% | 5/5 | 5/5 |
SoFi
5 original PDFs, 51 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 9.21 | 91.621% | 99.407% | 5/5 | 0/5 |
| Tesseract PSM 6 | 8.65 | 88.160% | 99.588% | 5/5 | 1/5 |
| OCRmyPDF | 14.22 | 90.528% | 98.807% | 4/5 | 1/5 |
| Docling native | 10.89 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 12.35 | 99.636% | 99.455% | 5/5 | 1/5 |
| Docling RapidOCR | 25.49 | 100.000% | 99.637% | 5/5 | 3/5 |
| Codex Sol | 18.88 | 99.636% | 99.636% | 5/5 | 4/5 |
Stockpile
1 original PDF, 8 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 1/1 | 1/1 |
| Tesseract PSM 3 | 25.35 | 100.000% | 100.000% | 1/1 | 1/1 |
| Tesseract PSM 6 | 22.25 | 100.000% | 100.000% | 1/1 | 1/1 |
| OCRmyPDF | 14.37 | 100.000% | 100.000% | 1/1 | 1/1 |
| Docling native | 7.77 | 100.000% | 100.000% | 1/1 | 1/1 |
| Docling Tesseract | 22.35 | 100.000% | 100.000% | 1/1 | 1/1 |
| Docling RapidOCR | 56.64 | 100.000% | 100.000% | 1/1 | 1/1 |
| Codex Sol | 35.67 | 100.000% | 100.000% | 1/1 | 1/1 |
Truist
5 original PDFs, 30 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 11.46 | 97.838% | 100.000% | 5/5 | 1/5 |
| Tesseract PSM 6 | 10.63 | 82.703% | 96.835% | 5/5 | 0/5 |
| OCRmyPDF | 8.79 | 98.378% | 100.000% | 5/5 | 2/5 |
| Docling native | 4.16 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling Tesseract | 9.08 | 98.919% | 100.000% | 5/5 | 3/5 |
| Docling RapidOCR | 25.32 | 100.000% | 100.000% | 5/5 | 5/5 |
| Codex Sol | 27.92 | 100.000% | 100.000% | 5/5 | 5/5 |
Wealthfront
5 original PDFs, 45 physical pages.
| Workflow | Median seconds | Raw recall | Raw precision | Passes | Exact PDFs |
|---|---|---|---|---|---|
| Embedded text | 0.08 | 100.000% | 100.000% | 5/5 | 5/5 |
| Tesseract PSM 3 | 5.95 | 99.874% | 100.000% | 5/5 | 4/5 |
| Tesseract PSM 6 | 5.46 | 99.874% | 99.874% | 5/5 | 4/5 |
| OCRmyPDF | 5.36 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling native | 3.76 | 99.748% | 99.748% | 5/5 | 3/5 |
| Docling Tesseract | 9.26 | 100.000% | 100.000% | 5/5 | 5/5 |
| Docling RapidOCR | 17.16 | 100.000% | 100.000% | 5/5 | 5/5 |
| Codex Sol | 17.61 | 100.000% | 100.000% | 5/5 | 5/5 |
What I’d use in Wombat
I’d keep native text and geometry as the first path, with checks for missing fields and image-only regions. For model-assisted recovery, I’d start with the tested gpt-6.1-sol medium configuration and retry flagged pages or rows at higher resolution. Its numeric recovery was useful on this sample; it doesn’t authorize ledger writes on its own.
I want the slow work to run as a resumable job: one original at a time, with its source hash, extraction settings and raw candidate preserved. Shared Rust should validate and publish financial facts. Unresolved values should remain explicit and correctable with provenance. A model shouldn’t invent missing basis.
This benchmark didn’t change the production extractor, ingest into the financial library or resolve the earlier parser findings. The next fixture work needs independently labeled financial rows and adversarial ownership errors, alongside the OCR disagreements. That’s where I can test whether the gain calculation is trustworthy.
Was this useful?
Thanks for the feedback.