Writing Data note

Technology · 26 min read

I benchmarked OCR because I didn't trust the numbers

I compared local OCR and Codex on full financial PDFs. The useful result was learning which checks an accuracy score leaves out.

In this article9 sections

Two Codex runs, configured with different model identifiers, agreed on the wrong price. Both returned structured output without flagging uncertainty.

I was testing extraction for Wombat because I wanted capital-gain and loss calculations I could trust. Changing the OCR engine could recover more numbers. I still needed to establish whether those numbers belonged to the right transactions.

I tested 78 distinct financial PDFs across 18 institution groups, covering all 531 physical pages. Fifteen groups had five originals each. Three had only one available original each; I didn’t count copies as extra samples. The corpus mixed statements, confirmations, cards, loans and retirement documents. These were mostly digital PDFs, sometimes with image-only regions. This says little about damaged scans or handwriting.

Benchmark numbers

Measured on an Apple M3 Mac with 16 GiB of RAM on October 9, 2026. One traversal per configuration; desktop activity and system swap weren’t isolated. These are observed workflow latencies, not repeated-run performance estimates. The scores use a source-derived reference corrected after examining disagreements; they do not measure verified capital-gain accuracy.

What the percentages mean

The unit is a decimal occurrence on a physical page. A repeated value counts again each time it appears. A match preserves the digits, printed decimal precision and the benchmark’s sign convention. These percentages do not measure dollars recovered or the correctness of a gain calculation.

MetricCalculationWhat it tells me
RecallMatched occurrences ÷ expected reference occurrences × 100How much of the reference did the extractor recover? Missing values lower recall.
PrecisionMatched occurrences ÷ emitted occurrences × 100How much of the extractor’s output matched the reference? Extra or wrong values lower precision.
Exact PDF ratePDFs with exact inventories and passing protocols ÷ attempted PDFs × 100How often did a whole PDF match on every page? One missing, extra or changed occurrence makes that PDF non-exact.

For an invented example, suppose the reference has 100 occurrences. The extractor emits 95, of which 90 match. Recall is 90 ÷ 100 = 90%. Precision is 90 ÷ 95 = 94.74%. Ten expected occurrences were not matched, and five emitted occurrences were not matched. A changed digit can hurt both scores: the expected value is missing and the wrong value is extra.

Matches are counted within each physical page, then pooled across the workflow. A document containing many decimal occurrences contributes more to recall and precision than a short document. The exact PDF rate gives each attempted PDF one vote. Even a perfect inventory can put basis on the wrong transaction row; this benchmark does not check row ownership.

Codex recovered 99.942% of expected decimal occurrences in 43.17 minutes, with exact inventories on 67 of 78 PDFs. Direct Tesseract PSM 3 took 14.89 minutes and had 38 exact PDFs. Embedded text took 0.10 minutes, but it supplied most of the reference and is not an independent accuracy result.

Horizontal bar chart of elapsed minutes for all eight workflows, from embedded text at 0.10 minutes to Codex Sol at 43.17 minutes.

The time bars include unsuccessful attempts. Rendering is included; Docling model warmup is excluded. These are single observed traversals on the same corpus.

Horizontal bar chart of exact numeric inventories out of 78 PDFs: embedded text 73, Docling text 65, Tesseract PSM 3 38, PSM 6 33, OCRmyPDF 39, Docling Tesseract 49, Docling RapidOCR 56, and Codex Sol 67.

An exact inventory requires the same decimal multiset on every physical page and a passing protocol. It does not check that the numbers belong to the right financial rows.

Scatter plot comparing elapsed minutes with the percentage of PDFs having exact inventories. Codex Sol is at 43.17 minutes and 85.9%; Docling native text is at 10.72 minutes and 83.3%. The embedded-text reference baseline is marked separately.

The scatter plot uses the exact PDF rate on its vertical axis. Codex’s point is 67 ÷ 78 × 100 = 85.90%, not its 99.942% decimal recall. Higher means more exact PDFs; farther left means less elapsed time. The table keeps decimal recall and precision separate from whole-PDF exactness.

WorkflowRaw decimal recallRaw decimal precisionProtocol passesExact PDFsAttempted minutes
Embedded PDF text*99.903%100.000%78/7873/780.10
Docling, native text99.749%99.807%78/7865/7810.72
Tesseract, PSM 398.388%99.677%78/7838/7814.89
Tesseract, PSM 696.775%99.514%78/7833/7813.64
OCRmyPDF, force OCR85.596%99.473%76/7839/7810.54
Docling, full-page Tesseract99.566%99.806%78/7849/7818.80
Docling, full-page RapidOCR99.759%99.518%78/7856/7838.83
Codex, gpt-6.1-sol / medium99.942%99.865%77/7867/7843.17

Embedded text isn’t OCR. Most of that row’s agreement is circular because it produced the text reference. The added source-image values expose known omissions in that layer.

For Codex, 10,352 of 10,358 expected occurrences matched, giving 99.942% recall. It emitted 10,366 decimal occurrences, giving 10,352 ÷ 10,366 = 99.865% precision. That leaves six unmatched expected occurrences and fourteen unmatched emitted occurrences under this checker. These are occurrence counts, not a count of bad transactions or incorrect tax calculations.

Protocol passes count PDFs whose pipeline completed and whose output met its required format. A protocol pass does not establish numerical or financial correctness. An exact PDF needs both a protocol pass and an exact inventory.

Raw recall and precision include valid decimal output from unsuccessful attempts. A failed pipeline or malformed response remains a failure; those values aren’t an accepted ingestion result. The Codex run completed its calls, but one PDF’s response included an integer in the required decimal-only array. JSON shape validation alone didn’t catch that contract error.

Tested versions and settings

These are the versions installed for the October 9, 2026 benchmark, not claims about the latest releases.

  • Machine: Apple M3, 16 GiB RAM; local CPU workflows.
  • Poppler 26.09.0: pdftotext -layout; embedded-text baseline, not OCR.
  • Tesseract 5.5.3: English, 300 dpi grayscale, PSM 3 and PSM 6 as separate runs; OMP_THREAD_LIMIT=2.
  • OCRmyPDF 17.13.0: force OCR, two jobs, PSM 3, optimization disabled; wraps Tesseract.
  • Docling 2.135.0: native-text, full-page Tesseract and full-page RapidOCR as three separate workflows; tables enabled, remote services disabled, roughly 216 dpi for full-page OCR.
  • RapidOCR 3.9.1: local PP-OCRv6 detection and recognition, plus the ch_ppocr_mobile_v2.0 orientation classifier.
  • ONNX Runtime 1.31.0: runtime for the RapidOCR models.
  • Codex CLI 0.162.0: existing ChatGPT login; 180 dpi color pages, at most four pages per chunk; 152 sequential calls for the full corpus.
  • Requested full-corpus model: gpt-6.1-sol, medium reasoning.
  • Requested comparison model: gpt-6-astra, medium reasoning, on eight originals only. The clean Sol comparison used gpt-6.1-sol / medium, with project instructions, memory, plugins and app tools disabled per invocation.

The Codex CLI did not expose a resolved backend model snapshot. The identifiers above describe what I requested.

What I measured

Every workflow attempted every selected original in full. I asked Codex to transcribe every visible decimal literal, preserving precision and repeated occurrences. It received page images, not the reference text. The local workflows produced text or table cells that I normalized into the same decimal inventory.

The checker compared a multiset per physical page. Missing values, extra values, changed digits and changed precision affected the score. Integers, dates without decimal points, ordinary text, reading order and ownership of values were outside the metric.

The reference came from the PDF’s embedded text, with ten values added after checking source images. It contained 10,358 decimal occurrences, including 1,349 zeros. Recall counts recovered expected occurrences; precision counts how many emitted occurrences match the reference. It was not an independently labeled human gold set. I corrected the derived checker after examining disagreements, so these are exploratory agreement figures. I haven’t manually audited all 531 pages. The PDFs and raw evidence remain private.

A parenthesis rule was also a convention: the prompt encoded parenthetic decimals as negatives, including explanatory disclosures where the financial meaning wasn’t negative. The private report separates magnitude-only recovery from that signed convention.

Codex timings include page rendering, CLI startup, network time, inference and output serialization. I recorded token usage, but this wasn’t a per-token API billing experiment. The event audit recorded no tool calls during transcription.

PSM 3 automatically segments a page; PSM 6 assumes a single block. Those layout assumptions are described in Tesseract’s quality guide. I charged rendering to each independently usable configuration.

OCRmyPDF wraps Tesseract and PDF processing. Docling adds layout and table reconstruction; its full-page OCR configuration is another workflow, not another independent recognizer. Its Tesseract path also performed orientation detection. Docling model warmup was logged separately. Two threads per configured model or stage did not mean a global two-core cap: pipelined stages could overlap.

Raster resolutions differed across workflows. I didn’t test every OCR product, GPU acceleration or arbitrary Codex model configurations.

I also kept timings separate by institution. A short card statement and a dense investment statement create different work; an overall average conceals that mix. The institution tables below include document counts, page counts, median whole-PDF latency, numeric agreement and failures. The selection favored available recent originals across document families; some templates were older. This was a purposive sample, not a random sample of financial PDFs.

The checker needed debugging too

Some apparent model errors were reference errors. A chart legend existed only as an image. Payment amounts were spread across character boxes, so the text layer didn’t contain a contiguous decimal. Dot leaders obscured printed minus signs. A date separator looked like a negative quantity to my regex.

I preserved the original text and candidate outputs, versioned the checker corrections, and applied each correction to every workflow. Those changes don’t make the benchmark blind or independent. They make the remaining comparison more useful.

A PDF can contain plenty of text and still omit visible numbers from its text layer. I can’t use “text exists” as proof that native extraction is complete.

The wrong numbers were plausible

The primary Codex run changed one purchase-price digit and misread a cash amount in two repeated fields. The repeated fields agreed with each other, so an equality check would have passed. Quantity multiplied by price caught the cash discrepancy, allowing for printed precision and cash rounding. That check concerns printed trade cash, not tax-adjusted basis; fees and source rounding need their own handling.

A separate eight-original comparison requested gpt-6-astra at medium reasoning. Both configurations had exact numeric inventories on the preselected five short originals. On three additional stress originals, the second configuration corrected the cash reading but repeated the wrong price. Those stress cases were selected after seeing a failure, so they aren’t unseen validation data. I can’t call model agreement an independent truth check or infer a global best model from this subset.

All three configurations used medium reasoning. The clean Sol configuration disabled project instructions, memory, plugins and app tools for each invocation. That changed the configuration, not just the requested model. These single traversals do not establish a causal speed improvement.

SubsetConfigurationExact PDFsSigned matchesTotal seconds
Five short originalsCodex Sol5/5160/16070.67
Five short originalsCodex Astra5/5160/16099.80
Five short originalsCodex Sol, clean config5/5160/16069.59
Three failure-selected originalsCodex Sol1/32448/2451286.56
Three failure-selected originalsCodex Astra2/32450/2451239.38
Three failure-selected originalsCodex Sol, clean config2/32449/2451243.22

The short subset covered SoFi, Wealthfront, Robinhood, Chase and Citi: 15 pages and 160 expected decimal occurrences. The additional stress subset covered SoFi, Wealthfront and Acorns: 31 pages and 2,451 occurrences. Every response in these subsets passed the protocol. The larger Astra comparison was not run across all 78 originals.

I retried the troublesome page three times at 180 dpi and three times at 300 dpi. The 180 dpi page passed twice; the 300 dpi page passed all three times. A 600 dpi crop of the source row also passed three times. This is evidence for a targeted higher-resolution retry, with a tiny sample and a changed single-page context. It isn’t a guarantee.

Retry inputExact inventories
Full page, 180 dpi2/3
Full page, 300 dpi3/3
Source row crop, 600 dpi3/3

The primary run reported no uncertainty entries anywhere, despite the confirmed mistakes. Confidence wasn’t a useful acceptance test here.

A perfect inventory can still produce wrong gains

This example uses invented transactions, not my account data:

ExtractionRowGiven categoryProceedsBasisGain
CorrectAlphaLong term12010020
CorrectBetaShort term806020
Basis rows swappedAlphaLong term1206060
Basis rows swappedBetaShort term80100-20

Both versions contain the same numbers, and both total 40 in gains. The category results differ. I ran this invented row swap through the actual inventory-matching function; it passed all four numeric occurrences. This demonstrates a limit of the benchmark, not a measured error from an OCR engine or a Wombat financial-parser test.

Tests for gains need to assert row ownership, account-specific lots, source basis, categories and final numbers. A numeric inventory or balanced total can’t replace those assertions.

Rendering can fail before recognition

The fixed OCRmyPDF workflow failed on two originals. One renderer couldn’t create a bitmap; another run produced an invalid output PDF. Both original PDFs passed a syntax check, which doesn’t establish financial correctness.

Switching to Ghostscript didn’t resolve either failure. One retry attempted a 1.5-billion-pixel image and hit its image-size guard. I kept the guard enabled.

A separate recovery pipeline rendered every original page at a bounded 300 dpi, rebuilt a derivative image PDF, then ran OCRmyPDF. Both recovered attempts produced valid PDFs with all physical pages retained. They still missed some decimal values. These were failure-led retries on two sources, not replacement scores for the main corpus.

The recovered 73-page brokerage PDF took 75.37 seconds and matched 1,130 of 1,142 expected decimals. The five-page bank PDF took 9.43 seconds and matched 23 of 26. Those times include the compound recovery workflow. Successful rendering still needed an extraction check.

Results by institution

Median seconds are per successfully completed whole PDF, including rendering; Docling model warmup is excluded. Each workflow attempted every selected PDF. Recall includes expected values from failed attempts. Precision includes any decimal output those attempts emitted; that partial output remains diagnostic. Exact means a matching decimal multiset on every physical page, with a passing protocol. It does not mean every financial field is correct. Percentages are rounded to three decimals.

Document types and lengths differ. These medians cannot rank how difficult a bank is to parse, and a median can hide one long document: Robinhood included a 73-page original. Apple / Goldman Sachs includes Marcus-branded documents. Rocket Mortgage is grouped by document branding. American Express evidence was available as a spreadsheet, without a dedicated PDF sample here.

Acorns

1 original PDF, 23 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.21100.000%100.000%1/11/1
Tesseract PSM 333.78100.000%100.000%1/11/1
Tesseract PSM 631.28100.000%100.000%1/11/1
OCRmyPDF23.0292.503%99.953%1/10/1
Docling native56.60100.000%100.000%1/11/1
Docling Tesseract71.61100.000%100.000%1/11/1
Docling RapidOCR136.47100.000%98.702%1/10/1
Codex Sol250.1299.956%99.956%1/10/1

Apple / Goldman Sachs

5 original PDFs, 17 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 35.97100.000%100.000%5/55/5
Tesseract PSM 65.49100.000%100.000%5/55/5
OCRmyPDF9.0399.351%99.351%5/54/5
Docling native2.05100.000%100.000%5/55/5
Docling Tesseract6.71100.000%100.000%5/55/5
Docling RapidOCR13.72100.000%99.355%5/54/5
Codex Sol14.08100.000%100.000%5/55/5

Capital One

5 original PDFs, 18 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.05100.000%100.000%5/55/5
Tesseract PSM 38.4295.270%99.296%5/52/5
Tesseract PSM 68.1091.216%100.000%5/51/5
OCRmyPDF5.7391.892%100.000%5/50/5
Docling native3.47100.000%100.000%5/55/5
Docling Tesseract9.1998.649%100.000%5/53/5
Docling RapidOCR17.13100.000%100.000%5/55/5
Codex Sol15.42100.000%98.013%5/52/5

Charles Schwab

5 original PDFs, 26 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.09100.000%100.000%5/55/5
Tesseract PSM 310.2498.415%98.763%5/52/5
Tesseract PSM 69.6198.592%98.765%5/53/5
OCRmyPDF10.6298.592%98.765%5/52/5
Docling native4.37100.000%99.475%5/54/5
Docling Tesseract12.75100.000%100.000%5/55/5
Docling RapidOCR30.0699.648%100.000%5/53/5
Codex Sol23.4499.472%98.949%5/52/5

Chase

5 original PDFs, 20 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 34.3799.401%100.000%5/54/5
Tesseract PSM 64.0598.802%98.214%5/53/5
OCRmyPDF4.24100.000%100.000%5/55/5
Docling native2.99100.000%100.000%5/55/5
Docling Tesseract6.12100.000%100.000%5/55/5
Docling RapidOCR12.23100.000%98.817%5/53/5
Codex Sol14.33100.000%100.000%5/55/5

Citi

5 original PDFs, 26 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 39.2298.413%100.000%5/53/5
Tesseract PSM 68.5596.032%100.000%5/52/5
OCRmyPDF6.5298.413%100.000%5/53/5
Docling native4.0399.206%100.000%5/54/5
Docling Tesseract11.1396.032%100.000%5/52/5
Docling RapidOCR21.9599.206%100.000%5/54/5
Codex Sol21.34100.000%100.000%5/55/5

Comenity

5 original PDFs, 24 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 39.3695.337%98.925%5/50/5
Tesseract PSM 68.6590.674%96.685%5/50/5
OCRmyPDF7.1895.855%97.884%5/50/5
Docling native3.11100.000%100.000%5/55/5
Docling Tesseract9.0398.964%99.479%5/54/5
Docling RapidOCR17.8798.964%100.000%5/53/5
Codex Sol21.11100.000%100.000%5/55/5

Discover

5 original PDFs, 18 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 34.6987.023%96.610%5/50/5
Tesseract PSM 64.5284.733%93.277%5/50/5
OCRmyPDF4.1196.947%98.450%5/53/5
Docling native3.28100.000%100.000%5/55/5
Docling Tesseract7.8997.710%98.462%5/53/5
Docling RapidOCR15.5899.237%100.000%5/54/5
Codex Sol15.98100.000%98.496%5/53/5

Fidelity

1 original PDF, 4 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.0496.685%100.000%1/10/1
Tesseract PSM 37.2899.448%100.000%1/10/1
Tesseract PSM 66.75100.000%100.000%1/11/1
OCRmyPDF5.4596.133%98.305%1/10/1
Docling native6.1496.685%100.000%1/10/1
Docling Tesseract11.05100.000%100.000%1/11/1
Docling RapidOCR24.6999.448%100.000%1/10/1
Codex Sol23.71100.000%100.000%1/11/1

J.P. Morgan

5 original PDFs, 80 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.13100.000%100.000%5/55/5
Tesseract PSM 339.8199.180%99.740%5/52/5
Tesseract PSM 636.7799.612%99.870%5/52/5
OCRmyPDF26.1197.454%99.122%5/52/5
Docling native38.52100.000%100.000%5/55/5
Docling Tesseract51.1599.741%99.957%5/52/5
Docling RapidOCR103.16100.000%100.000%5/55/5
Codex Sol63.17100.000%100.000%4/54/5

Morgan Stanley

5 original PDFs, 26 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 35.6699.627%99.627%5/53/5
Tesseract PSM 65.4293.657%99.735%5/53/5
OCRmyPDF5.3799.627%99.627%5/53/5
Docling native1.8099.627%100.000%5/54/5
Docling Tesseract8.1299.254%100.000%5/51/5
Docling RapidOCR14.4799.502%100.000%5/53/5
Codex Sol18.83100.000%100.000%5/55/5

Robinhood

5 original PDFs, 87 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 38.5897.617%99.620%5/54/5
Tesseract PSM 68.0293.150%99.365%5/50/5
OCRmyPDF6.2714.520%100.000%4/52/5
Docling native5.1199.255%99.701%5/52/5
Docling Tesseract13.5299.181%99.701%5/51/5
Docling RapidOCR22.8399.255%99.701%5/52/5
Codex Sol18.00100.000%100.000%5/55/5

Rocket Mortgage

5 original PDFs, 10 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.0497.647%100.000%5/51/5
Tesseract PSM 34.2990.588%98.718%5/51/5
Tesseract PSM 64.1289.412%97.436%5/51/5
OCRmyPDF3.9191.765%100.000%5/51/5
Docling native3.1997.647%93.785%5/51/5
Docling Tesseract5.9296.471%94.798%5/51/5
Docling RapidOCR8.5797.647%93.785%5/51/5
Codex Sol17.96100.000%100.000%5/55/5

ScholarShare 529

5 original PDFs, 18 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.04100.000%100.000%5/55/5
Tesseract PSM 34.90100.000%100.000%5/55/5
Tesseract PSM 64.57100.000%100.000%5/55/5
OCRmyPDF3.66100.000%100.000%5/55/5
Docling native2.65100.000%100.000%5/55/5
Docling Tesseract5.93100.000%100.000%5/55/5
Docling RapidOCR9.75100.000%100.000%5/55/5
Codex Sol12.29100.000%100.000%5/55/5

SoFi

5 original PDFs, 51 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 39.2191.621%99.407%5/50/5
Tesseract PSM 68.6588.160%99.588%5/51/5
OCRmyPDF14.2290.528%98.807%4/51/5
Docling native10.89100.000%100.000%5/55/5
Docling Tesseract12.3599.636%99.455%5/51/5
Docling RapidOCR25.49100.000%99.637%5/53/5
Codex Sol18.8899.636%99.636%5/54/5

Stockpile

1 original PDF, 8 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%1/11/1
Tesseract PSM 325.35100.000%100.000%1/11/1
Tesseract PSM 622.25100.000%100.000%1/11/1
OCRmyPDF14.37100.000%100.000%1/11/1
Docling native7.77100.000%100.000%1/11/1
Docling Tesseract22.35100.000%100.000%1/11/1
Docling RapidOCR56.64100.000%100.000%1/11/1
Codex Sol35.67100.000%100.000%1/11/1

Truist

5 original PDFs, 30 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 311.4697.838%100.000%5/51/5
Tesseract PSM 610.6382.703%96.835%5/50/5
OCRmyPDF8.7998.378%100.000%5/52/5
Docling native4.16100.000%100.000%5/55/5
Docling Tesseract9.0898.919%100.000%5/53/5
Docling RapidOCR25.32100.000%100.000%5/55/5
Codex Sol27.92100.000%100.000%5/55/5

Wealthfront

5 original PDFs, 45 physical pages.

WorkflowMedian secondsRaw recallRaw precisionPassesExact PDFs
Embedded text0.08100.000%100.000%5/55/5
Tesseract PSM 35.9599.874%100.000%5/54/5
Tesseract PSM 65.4699.874%99.874%5/54/5
OCRmyPDF5.36100.000%100.000%5/55/5
Docling native3.7699.748%99.748%5/53/5
Docling Tesseract9.26100.000%100.000%5/55/5
Docling RapidOCR17.16100.000%100.000%5/55/5
Codex Sol17.61100.000%100.000%5/55/5

What I’d use in Wombat

I’d keep native text and geometry as the first path, with checks for missing fields and image-only regions. For model-assisted recovery, I’d start with the tested gpt-6.1-sol medium configuration and retry flagged pages or rows at higher resolution. Its numeric recovery was useful on this sample; it doesn’t authorize ledger writes on its own.

I want the slow work to run as a resumable job: one original at a time, with its source hash, extraction settings and raw candidate preserved. Shared Rust should validate and publish financial facts. Unresolved values should remain explicit and correctable with provenance. A model shouldn’t invent missing basis.

This benchmark didn’t change the production extractor, ingest into the financial library or resolve the earlier parser findings. The next fixture work needs independently labeled financial rows and adversarial ownership errors, alongside the OCR disagreements. That’s where I can test whether the gain calculation is trustworthy.

Was this useful?

What was missing?

Thanks for the feedback.