LiteParse September update: PDFium 20-25% faster, plus visual grounding and is-complex routing
What happened
LiteParse 这次更新主要干了四件事。第一,他们给底层的 PDFium 引擎动了次精细手术,把文本提取的耗时砍掉了 20-25%。不开 OCR 的情况下,纯文本提取平均每页只要 2.8 毫秒,完整转成 markdown 是 3.9 毫秒,这是他们测过的开源解析器里最快的。第二,markdown 的启发式转换规则更准了,但正文没给出具体的准确率数字...
Coverage
Follow the reports to see the story from different sides.
- AI HOT (Curated Pool)LiteParse September update: PDFium 20-25% faster, plus visual grounding and is-complex routing
LiteParse shipped four updates. First, a fork of PDFium with surgical optimizations cuts text extraction time by 20-25%. With OCR off, it averages 2.8ms/page for text and 3.9ms/page for full markdown rendering—the fastest open parser they've tested. Second, markdown heuristics accuracy improved, though the post doesn't share specific metrics. Third, visual grounding now maps parsed elements back to PDF page coordinates. Fourth, a new is-complex API lets callers route documents by complexity before choosing a parsing pipeline. LiteParse currently sees 300k+ weekly downloads and 12k+ GitHub stars.