LlamaIndex introduces just-in-time agentic OCR: a two-pass document processing method that balances cost and accuracy
LlamaIndex 解析 just-in-time Agentic OCR:两遍式文档处理如何平衡成本与精度
LlamaIndex splits document processing into two passes: a fast first pass with lightweight OCR to extract text, and a second pass that calls a vision model (VLM) only when needed for charts or scanned pages. They tested this on 84 SEC filings, using metadata and text retrieval to find relevant pages before running deep OCR on just a few. This works for interactive Q&A in a data room but not for offline batch pipelines—each query can trigger new VLM calls, so latency and cost grow with the number of questions. The post does not disclose specific cost comparisons or latency figures.
Why it matters: LlamaIndex's two-pass document processing is a solid engineering piece with real numbers on 84 SEC filings and honest scope limits. But it's a toolchain optimization, not a model or product launch, so industry impact is modest. 72 lands right at the featured threshold.