Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Microsoft blames massive Microsoft 365 outage on maintenance bug

    July 25, 2026

    Samsung Wallet Will Add Stablecoin Support, Including USDC

    July 25, 2026

    The Economic Philosophy of Britain’s Andy Burnham

    July 25, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Microsoft blames massive Microsoft 365 outage on maintenance bug
    • Samsung Wallet Will Add Stablecoin Support, Including USDC
    • The Economic Philosophy of Britain’s Andy Burnham
    • I grew up near Andy Burnham. This is what shaped our new PM | Andy Burnham
    • The UK’s first TikTok PM? Andy Burnham channels Zohran Mamdani as he hits social media | Labour
    • India and South Africa lead push to amass emergency fuel stockpiles
    • Samsung Galaxy Watch 9 vs. Google Pixel Watch 4: I compared both Android flagships, here’s what I prefer
    • Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Saturday, July 25
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKJuly 25, 2026 Artificial Intelligence No Comments7 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Datalab has released Marker 2, a full rewrite of its open source document conversion pipeline. Marker converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files into markdown, JSON, HTML, or chunks. The Datalab team rebuilt it around three components shipped over the preceding months: Surya OCR 2, a 20M-param fast layout model, and a rebuilt pdftext that is 3Ă— faster than the previous one.

    The main result comes from olmOCR-bench, a third-party benchmark from Allen AI. Marker 2’s balanced mode scores 76.0% overall and 83.5% on born-digital PDFs. It sustains 2.9 pages per second on a single B200 GPU. That is over 5× the throughput of MinerU’s pipeline backend, which scores 72.7% at 0.54 pages per second. Docling scores 50.3% at 2.1 pages per second on the same harness.

    What’s New in Marker 2

    Marker 2 exposes three conversion paths instead of one:

    • balanced — the Surya VLM handles layout, and the whole page is re-OCR’d whenever embedded text is bad. Highest quality, best on GPU. 76.0% olmOCR-bench.
    • fast — a lightweight rf-detr/onnx layout detector plus pdftext, with minimal, surgical VLM use. 66.6%, and far cheaper.
    • –disable_ocr — pure text-layer extraction, no VLM calls at all. Runs entirely on CPU. 43.6%, 23.7 pg/s.

    Mode is now device-aware by default: balanced on GPU, fast on CPU/MPS, overridable with –mode. Full CPU support is the second structural change. fast –disable_ocr needs no GPU and no inference server, and the 20M layout model still reads columns, tables and headers on CPU.

    The third change is architectural, and it is the one that produces the throughput numbers. Many thin CPU workers share a single Surya inference server. The parent process budgets VLM concurrency across them, so throughput scales with server capacity rather than per-process VRAM. Datalab reports that balanced mode sustains ~2.9 pg/s against a ~0.3 pg/s single-stream rate on the same hardware.

    Breaking changes are worth flagging before an upgrade. Python 3.10+ is now required. Packaging moved from Poetry to uv, with hatchling as the build backend, though pip install marker-pdf is unchanged. The structured-extraction converter and extractors were removed; Datalab points users to the hosted API or a –use_llm workflow instead.

    Comparison

    The scoring benchmark is olmOCR-bench from Ai2: 1,403 PDFs with roughly 8,400 pass/fail unit tests covering math rendering, table structure, reading order, headers and footers, and old scans. The overall score is the macro-average across the 8 categories, computed with the official olmOCR-bench checker. Throughput is sustained concurrent pg/s on one B200 host, not single-stream latency.

    A note on provenance. olmOCR-bench is a third-party benchmark from Ai2, but every score and throughput figure below comes from Datalab’s own runs. All of them are reproducible through the open harness in the Marker repository, which ships competitor runners for MinerU, Docling and LiteParse alongside Marker’s own.

    These numbers also reflect one benchmark’s document mix measured on a single hardware setup, so results on your own documents may differ. Teams evaluating these systems should run the harness against their own corpus, which is the only way to know how the four rank on the documents they actually process.

    Marker 2 vs MinerU

    MinerU’s pipeline backend is the closest architectural match. Both read the PDF text layer and OCR selectively. On overall score, Marker balanced leads 76.0 to 72.7. On born-digital documents the two are effectively tied: 83.5 against 83.3.

    The separation is throughput. Marker balanced sustains 2.9 pg/s against MinerU’s 0.54 pg/s, a 5.4× gap at a higher score. Marker fast sustains 7.4 pg/s, roughly 13.7× MinerU’s pipeline rate, but scores 6.1 points below MinerU to do it.

    MinerU also ships a VLM backend, which Datalab states scores higher than its pipeline backend. That backend is a full-page-VLM approach and is not in this table. AI teams evaluating MinerU should benchmark that path separately.

    Marker 2 vs Docling

    Docling is the widest margin among the GPU pipelines. Marker balanced leads 76.0 to 50.3 overall and 83.5 to 64.0 on born-digital, while also running faster: 2.9 pg/s against 2.1 pg/s. Datalab notes Docling was run on its default pipeline, which uses the text layer for born-digital pages and OCR for image regions.

    Docling’s counterweight is governance and format breadth, not accuracy. The codebase is MIT-licensed, it originated at IBM Research, and it is hosted as a project in the LF AI & Data Foundation. Its input list also extends past documents into audio and email formats.

    Marker 2 vs LiteParse

    LiteParse, from the LlamaIndex team, is a Rust document parser. It does not compete on the same axis. On CPU it scores 22.4 overall and 20.4 with OCR off, against Marker’s CPU-only 43.6.

    But LiteParse with OCR disabled reports 1721 pg/s — roughly 73× Marker’s CPU mode, which is the tradeoff. Marker’s fast –disable_ocr runs a 20M layout model on CPU and still recovers structure, which is why it more than doubles a plain text dump’s score. LiteParse has no layout model and collapses on anything non-linear.

    Marker 2 vs the full-page VLM tier

    The Datalab team emphasizes that Marker is designed as a pipeline rather than a VLM, clarifying that these are distinct tools. In this evaluation, their hosted Chandra 2 scores 85.8, while Gemini Flash 3.5 via API scores 76.4. Datalab’s Chandra repository also positions Ai2’s olmOCR 2 at 82.4 and dots.ocr 1.5 at 83.9 within a separate table. For scans, math-heavy pages, and achieving top accuracy, the VLM tier remains superior to all listed pipelines.

    Marker’s balanced mode narrows the performance gap to just 0.4 points behind Gemini Flash 3.5 overall, and it even outperforms it on born-digital documents by a margin of 83.5 to 79.1—without requiring a per-page API call.

    Per-category behavior

    The mode you pick changes the failure profile, not just the score. Each row is one olmOCR-bench category, scored across all three modes. Math is the sharp edge: fast mode reads equations from the PDF text layer instead of VLM-OCRing them, so arXiv math falls from 83.9 to 23.4, and –disable_ocr scores 0.0 there by design. Outside the two math categories, old scans is the weakest split in every mode, topping out at 43.2.

    Licensing

    This is where the four systems diverge most for commercial teams:

    • Marker: code is Apache 2.0. Model weights use a modified AI Pubs OpenRAIL-M license — free for research, personal use, and startups under $5M funding/revenue. Beyond that, commercial use of the weights requires a paid license.
    • MinerU: now under the MinerU Open Source License, based on Apache 2.0 with added conditions. A separate commercial license is required above 100M MAU or $20M monthly revenue, and online services built on it must disclose that fact.
    • Docling: MIT, with model licenses tracked separately in their original packages.
    • LiteParse: open source, from run-llama, with LlamaParse positioned as the paid cloud path for hard documents.

    Use Case- Comparison

    Score alone does not pick the tool. Corpus type, hardware, licensing band and output format decide it. Try the interactive picker below to filter ten deployment scenarios by constraint and by tool, and see which parser fits your use case.

    Key Takeaways

    • Marker 2 balanced scores 76.0% on olmOCR-bench at 2.9 pg/s — over 5Ă— MinerU’s pipeline throughput.
    • It beats Docling on both axes at once: 76.0% against 50.3%, and 2.9 pg/s against 2.1 pg/s.
    • LiteParse trades structure for speed — 1721 pg/s with OCR off, but 20.4% against Marker’s 43.6% on CPU.
    • Fast mode with –disable_ocr runs entirely on CPU, no inference server, at 23.7 pg/s.
    • Licensing splits the field: Docling is MIT, MinerU stays free to $20M monthly revenue, and Marker’s weights need a paid license above $5M.
    • All benchmark and throughput numbers ship with a reproducible benchmarks/ harness.

    Interactive Dynamic Explainer


    Links: GitHub repo | Release notes | Blog post | Announcement tweet


    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

    Benchmark Breakdown Datalab Docling Liteparse Marker MinerU
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing

    Meta, Microsoft, Nvidia, IBM, and others back open-weight AI

    OpenAI pushes ChatGPT into patient health records

    OpenAI Presence: enterprise AI agents, engineers included

    How to Build an End-to-End OCR Pipeline with Baidu’s Unlimited-OCR for High-Resolution Images and Multi-Page PDF Parsing

    You Didn’t Get the AI Model You Paid For

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Microsoft blames massive Microsoft 365 outage on maintenance bug

    July 25, 2026

    Samsung Wallet Will Add Stablecoin Support, Including USDC

    July 25, 2026

    The Economic Philosophy of Britain’s Andy Burnham

    July 25, 2026

    I grew up near Andy Burnham. This is what shaped our new PM | Andy Burnham

    July 25, 2026
    Latest Posts

    Trump slaps 50% tariffs on Canada and Carney vows to ‘intensify’ trade talks

    July 21, 2026

    How Two Brothers Dug for Dead Relatives: With a Shovel and a Kitchen Knife

    July 21, 2026

    Chile floods: Towns evacuated following heavy rain in Coquimbo

    July 21, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Microsoft blames massive Microsoft 365 outage on maintenance bug

    July 25, 2026

    Samsung Wallet Will Add Stablecoin Support, Including USDC

    July 25, 2026

    The Economic Philosophy of Britain’s Andy Burnham

    July 25, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.