Skip to content
AI HOT (Curated Pool)

Tencent Hunyuan open-sources HyOCR-1.5: a 1B end-to-end OCR model with 6.37× faster inference

腾讯混元发布 HyOCR-1.5:端到端 OCR 大模型全栈开源,推理提速 6.37 倍

Tencent Hunyuan fully open-sourced HyOCR-1.5—training, inference, and model weights—a first for end-to-end OCR large models. The 1B-parameter model handles 8+ text-centric tasks and scores 94.74 on OmniDocBench v1.6, ranking first end-to-end. DFlash speculative decoding speeds up inference 6.37× under Transformers and 2.14× under vLLM, hitting 1.408s per page. It supports 4K resolution and a 128K context window, and uses Agentic Data Flow to extend low-resource OCR to 331 languages, ancient script recognition, and multi-image QA.

Why it matters: Tencent Hunyuan fully open-sourced an end-to-end OCR model — training code, inference code, weights. 1B params, 94.74 on OmniDocBench v1.6 (#1 among end-to-end models), 6.37x inference speedup via DFlash. A genuine open-source move from a major Chinese lab, not weights-only. S...

Read the original ↗Export Markdown