Tencent Hunyuan releases 1-bit and 4-bit quantized Hy3, a 295B MoE that runs on a single GPU
Tencent Hunyuan quantized its flagship Hy3 (295B MoE) into 1-bit and 4-bit versions that run on a single GPU. Hy3 is claimed to be best-in-class at this scale and competitive with trillion-parameter models for most agent scenarios. The quantized versions work via llama.cpp with MTP support, drastically lowering hardware requirements. Apache 2.0 license, commercial use allowed, plus two weeks of free API through OpenRouter. The post doesn't disclose quantization accuracy loss or the specific GPU memory needed.
Why it matters: Tencent Hunyuan's quantized Hy3 puts a 295B MoE model on a single GPU — immediately actionable for local deployment and agent builders. Apache 2.0 license plus a two-week free API window lowers the barrier to test. Held below 85 because the post doesn't disclose quantization a...