backburner explores iPhone and Mac collaborative inference
What happened
On October 10, Computing Life published a code walkthrough of the open-source project backburner, which lets an iPhone share local large-model inference with a Mac over USB-C. The iPhone handles prefill for long text and hosts the overflowed historical KV cache. The post covers the project's attention splitting, pipeline design, and how historical key-value pages are compiled into static ANE weights. In a 140k-context test, prefill throughput rose from 58 to 68 tok/s.
Written by AI from the coverage · updated 1 hour ago
Coverage
Follow the reports to see the story from different sides.
- Computing Life · Share · Yagebackburner 如何用 iPhone 协同 Mac 加速本地大模型
开源项目 backburner 让 iPhone 通过 USB-C 协同 Mac 运行本地大模型,分担长文本预填计算并托管溢出的历史 KV cache。文章基于代码解读其注意力切分、流水线和将历史键值页编译为 ANE 静态权重的方案,140k 上下文测试中预填吞吐从 58 提升至 68 tok/s。
Heat over time
Not enough continuous observations to draw a trend yet.