Trending storyPast story
Hugging Face transformers now runs GGUF quantized models directly
1 report1 sourceupdated 8 days ago
What happened
Summary
transformers 开始原生支持 GGUF 文件,你可以直接用 from_pretrained 加载一个 GGUF 模型,在本地跑起来。GGUF 是 llama.cpp 社区常用的模型压缩格式,能把模型体积压到适合笔记本内存的大小,代价是牺牲一点精度。这次更新背后复用了 llama.cpp 的 ggml 底层计算核心,优先在苹果芯片上做了优化,所...
Coverage
Follow the reports to see the story from different sides.
Sep 22
- AI HOT (Curated Pool)PickHugging Face transformers now runs GGUF quantized models directly
transformers now loads GGUF files natively, with local inference speed close to llama.cpp. You can use from_pretrained to load a GGUF checkpoint and run models like Qwen3.5 on a Mac. It reuses llama.cpp's ggml kernels under the hood, with initial optimization targeting Apple Silicon. Only the Qwen3.5 architecture is supported for now; more models and features are coming.