AI News Feed
Market watch
Products & Applications

Datalab Releases OmniExtractBench, an Open Benchmark for Structured Document Extraction

Datalab launched OmniExtractBench, an open benchmark for PDF-to-JSON extraction, with 620 documents and an explainable scorer.

The benchmark pools 620 documents from four existing benchmarks: 329 from LlamaIndex's ExtractBench, 202 from Datalab's own synthetic suite, 47 from micro1's LongExtractBench, which was commissioned by Reducto, and 42 from Extend's LongArray-Extract. Regulatory filing forms are the largest category at 88 documents. 128 documents are a single page, while 33 documents over 100 pages hold 40% of all pages. Datalab's synthetic suite is the second largest share.

OmniExtractBench gives each system a PDF and a JSON schema, and the system returns JSON that is scored value by value against a gold file. The scorer is available from PyPI as omni-extract-bench v0.1.7 for Python 3.11 and later, requires only SciPy, and is licensed under Apache 2.0. The code is on GitHub, and the data is on Hugging Face under CC BY 4.0. Rerunning vendors requires a user's own API keys and paid credits. Datalab said the launch post names four recurring problems with existing extraction benchmarks: bias toward the vendor that built the benchmark, opaque harnesses that can make a low score reflect a broken harness rather than a weak model, unclear scoring that leaves readers unable to tell why a document scored low, and narrow document variety.

The scorer flattens predicted and gold JSON into addresses, which are paths to single values, and normalizes each value first so that '03/31/2024' matches '2024-03-31.' Tables are the hard part. Compared by position, one missed row shifts every row after it; Datalab said it reran the scorer on a 100-row table missing its first row, where positional comparison scored 0% while OmniExtractBench scored 99%. The fix is content-based pairing with the Hungarian algorithm. ExtractBench and LongArray-Extract already align rows this way, Datalab said, and OmniExtractBench adds a verdict layer on top.

The verdict layer assigns one of six verdicts per value. Matched means paired and the values agree; misread means paired but the values differ; unfound means gold has a value and the prediction does not; fabricated means the schema allows it, gold is silent, and the prediction fills it; invented_item means part of a predicted row pairs with nothing; and invented_field means an address the schema never declared. Accuracy is matched values over all verdicts, precision divides matched values by predicted values, and recall divides them by gold values. The full rules are in the metric spec.

The null rule counts empty strings, None and whitespace as omissions, so those addresses are dropped. Strings such as 'NA' or '-' remain real answers. Datalab said this blocks a quiet exploit: padding a schema with empty optional fields to earn free matches. In its test, padded null fields added 0 verdicts.

Compared with other extraction benchmarks, OmniExtractBench has 620 documents from four sources, uses Hungarian content-based row alignment, offers per-value explanations with six verdict types, and carries an Apache 2.0 scorer license and CC BY 4.0 data license. ExtractBench by LlamaIndex has 370 documents, Hungarian alignment, per-field diffs and an HTML report, and Apache 2.0 licenses for scorer and data. LongArray-Extract by Extend has 45 synthetic documents, Hungarian alignment, per-document scores, an unstated scorer license, and CC BY 4.0 data. LongExtractBench by micro1 has 225 documents, 50 public, by-row-key alignment, an unstated scorer license, and MIT licensing with CC BY 4.0 labels only. The comparison was checked on Sept. 27, 2026.

Datalab scored 10 system configurations on the full corpus. Its accurate mode led at 93.85 accuracy. Datalab balanced at 93.48 and Reducto deep_extract v2 at 93.47 were effectively tied. Precision and recall then show how each system fails. Datalab in both modes and Reducto keep precision and recall within 0.6 points. GPT 5.6-sol posts 95.11 precision but 84.99 recall, losing 11.88% to unfound values; Gemini and Claude show the same pattern less sharply. LlamaExtract has 93.13 recall but 86.57 precision, losing 9.03% to fabricated values, while Extend loses 4.01% to invented items. Mistral OCR 4.1 and Azure Content Understanding trail on both metrics, with recall lower still.

Datalab provides a command to run the scorer: uv pip install omni-extract-bench, followed by oeb score --pred pred.json --gt gold.json --schema schema.json --verdicts. Installing the benchmark extra and running oeb benchmark reruns vendors, though users need their own API keys and paid credits.