Hugging Face transformers now runs GGUF quantized models directly
transformers 支持直接加载 GGUF 量化模型,本地推理性能接近 llama.cpp
transformers now loads GGUF files natively, with local inference speed close to llama.cpp. You can use from_pretrained to load a GGUF checkpoint and run models like Qwen3.5 on a Mac. It reuses llama.cpp's ggml kernels under the hood, with initial optimization targeting Apple Silicon. Only the Qwen3.5 architecture is supported for now; more models and features are coming.
Why it matters: HuggingFace adding native GGUF support to transformers bridges the most popular quantization format with the mainstream library, lowering the local-inference bar again. Score stays at 78 rather than higher because this is ecosystem plumbing, not a new capability breakthrough, ...