Skip to content
AI HOT (Curated Pool)

Tencent Hunyuan open-sources AngelSpec, a speculative decoding framework that more than doubles inference speed

腾讯混元开源 AngelSpec 投机解码框架

Tencent Hunyuan released AngelSpec, an end-to-end speculative decoding framework with training and deployment code. On the Hy3-A21B model, its DFly method achieves 1.98–2.40× end-to-end speedup over autoregressive decoding, and 10.5–11.8% higher throughput than DFlash. Draft model weights for Hy3-A21B MTP/DFly are also open-sourced.

Why it matters: Tencent Hunyuan open-sourced AngelSpec, a speculative decoding framework with full training and deployment code, showing 1.98-2.4x end-to-end speedup on Hy3-A21B and 10%+ throughput gains over DFlash. Speculative decoding is a hot inference optimization area, and a complete op...

Read the original ↗Export Markdown