Skip to content
Trending storyPast story

Hugging Face transformers now runs GGUF quantized models directly

1 report1 sourceupdated 8 days ago

What happened

Summary

transformers 开始原生支持 GGUF 文件,你可以直接用 from_pretrained 加载一个 GGUF 模型,在本地跑起来。GGUF 是 llama.cpp 社区常用的模型压缩格式,能把模型体积压到适合笔记本内存的大小,代价是牺牲一点精度。这次更新背后复用了 llama.cpp 的 ggml 底层计算核心,优先在苹果芯片上做了优化,所...

Coverage

Follow the reports to see the story from different sides.

Sep 22
  1. AI HOT (Curated Pool)Pick
    Hugging Face transformers now runs GGUF quantized models directly

    transformers now loads GGUF files natively, with local inference speed close to llama.cpp. You can use from_pretrained to load a GGUF checkpoint and run models like Qwen3.5 on a Mac. It reuses llama.cpp's ggml kernels under the hood, with initial optimization targeting Apple Silicon. Only the Qwen3.5 architecture is supported for now; more models and features are coming.