Skip to content
r/LocalLLaMA

PaddlePaddle releases HPD-Parsing: a 1B model hits 4,752 TPS for document parsing, 1.62× faster than the previous fastest parser

PaddlePaddle/HPD-Parsing · Hugging Face

PaddlePaddle released HPD-Parsing on Hugging Face, a 1B-param document parsing model. It uses a main layout branch for global coordination and dispatches localized content to parallel branches, with progressive multi-token prediction cutting decoding steps further. On OmniDocBench v1.6 it scores 94.91% overall—a new SOTA among end-to-end unified parsers—and peaks at 4,752 TPS, 1.62× the previous fastest parser and 3.06× its own autoregressive baseline. Training uses staged adaptation with automated difficulty-aware data curation to preserve accuracy. The post doesn't disclose hardware specs or VRAM requirements, so real-world cost needs your own testing.

Why it matters: PaddlePaddle drops a 1B doc parsing model that replaces token-by-token generation with hierarchical parallel decoding — clear architectural novelty, directly relevant to local RAG and doc processing practitioners. Missing benchmarks and concrete latency numbers, so 72 for now.

Read the original ↗Export Markdown