Skip to content
r/LocalLLaMA

Vision-capable LLMs vs. OCR for long-document QA with charts, images, and tables

Vision-capable LLMs vs. OCR for long-document (including charts, images, tables, etc.) QA

The author tested Claude Sonnet 4.5 on 171 questions from 30 image-heavy MMLongBench-Doc PDFs, comparing native PDF vision use with OCR pipelines. Native PDF ranked fifth of six at 52.0% accuracy and cost $0.2552 per query, while LlamaCloud premium with full context reached 59.6% at $0.1885 per query.

Why it matters: HKR-H/K/R pass: the post gives 30 PDFs, 171 questions, accuracy, and per-question cost for long-document QA. Limited sample and Reddit sourcing keep it in the featured-threshold band.

Read the original ↗Export Markdown