Skip to content
AI HOT (Curated Pool)

Tencent Hunyuan open-sources production inference kernels and integrates them into SGLang

HPC-Ops × SGLang:腾讯混元开源高性能 Attention、Router GEMM 与 MoE 算子

Tencent Hunyuan open-sourced its production-proven Attention, Router GEMM, and MoE kernels as HPC-Ops and merged them into SGLang's main branch. On H20, Hy3-FP8 TPOT dropped by up to 48.8%; the Attention kernel averaged 2.25× faster than FlashInfer/FlashAttention, and Router GEMM ran 1.30–3.22× faster than FP32 cuBLAS with lower error. H200 validation passed, with the MoE kernel hitting 4.21× over Triton on Qwen3.

Why it matters: Tencent Hunyuan open-sourced production-hardened Attention, Router GEMM, and MoE kernels and merged them into SGLang main. Three concrete optimizations with reproducible detail — not vague 'we made it faster' claims. High practical value for inference engineers, but audience i...

Read the original ↗Export Markdown