UnslothAI Releases Qwen3.6 MTP GGUF Models With Over 1.4x Faster Inference
UnslothAI发布Qwen3.6 MTP GGUF模型,实现推理速度大幅提升
Daniel Han released experimental Qwen3.6 MTP GGUF models, with the 27B model reaching 140 tokens/s on one GPU and the 35B-A3B version reaching 220 tokens/s, using two draft tokens for speculative decoding.
Why it matters: HKR-H/K/R pass via concrete single-GPU speed claims and local-inference relevance. Score stays in low featured because the post is a single X source and does not disclose GPU, quantization settings, or repro steps.