Skip to content
AI HOT (Curated Pool)

Tencent Hunyuan releases 1-bit and 4-bit Hy3, a 295B model that runs on a single GPU

腾讯混元发布 Hy3 1-bit 与 4-bit 版本

Tencent Hunyuan released 1-bit and 4-bit quantized versions of Hy3, a 295B model that runs on a single GPU via llama.cpp with MTP. The post doesn't disclose benchmarks or hardware requirements, but GGUF files are already on HuggingFace.

Why it matters: Tencent Hunyuan compressed its 295B model into 1-bit and 4-bit variants, using llama.cpp + MTP inference with GGUF downloads available. The technical path is concrete and the compression ratio is eye-catching, but no accuracy loss or hardware benchmarks are disclosed, so real-...

Read the original ↗Export Markdown