Skip to content
AI HOT (Curated Pool)

Hugging Face transformers now runs GGUF quantized models directly

transformers 支持直接加载 GGUF 量化模型,本地推理性能接近 llama.cpp

transformers now loads GGUF files natively, with local inference speed close to llama.cpp. You can use from_pretrained to load a GGUF checkpoint and run models like Qwen3.5 on a Mac. It reuses llama.cpp's ggml kernels under the hood, with initial optimization targeting Apple Silicon. Only the Qwen3.5 architecture is supported for now; more models and features are coming.

Why it matters: HuggingFace adding native GGUF support to transformers bridges the most popular quantization format with the mainstream library, lowering the local-inference bar again. Score stays at 78 rather than higher because this is ecosystem plumbing, not a new capability breakthrough, ...

Read the original ↗Export Markdown