Article may be outdated

This article is 67 days old. Some details may have changed since publication.

MarkTechPost·5 min read·hard

Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown

A
Asif Razzaq
Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown
✦AI Summary

Datalab has released Marker 2, an open-source document conversion pipeline that significantly outperforms competitors like MinerU and Docling in speed and accuracy. The update introduces device-aware modes and architectural changes to improve throughput for processing various document formats into structured data.

Why it matters

Efficient document conversion is critical for training large language models, and these benchmarks provide a standard for evaluating data ingestion pipelines.

✦Dive DeeperCreate a free account to unlock

Datalab has released Marker 2 , a full rewrite of its open source document conversion pipeline. Marker converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files into markdown, JSON, HTML, or chunks. The Datalab team rebuilt it around three components shipped over the preceding months: Surya OCR 2 , a 20M-param fast layout model , and a rebuilt pdftext that is 3× faster than the previous one.

The main result comes from olmOCR-bench , a third-party benchmark from Allen AI. Marker 2’s balanced mode scores 76.0% overall and 83.5% on born-digital PDFs. It sustains 2.9 pages per second on a single B200 GPU. That is over 5× the throughput of MinerU’s pipeline backend, which scores 72.7% at 0.54 pages per second. Docling scores 50.3% at 2.1 pages per second on the same harness.

Marker 2 exposes three conversion paths instead of one:

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in