Skip to content
r/LocalLLaMA

Paper on Hummingbird+: low-cost FPGAs for LLM inference

[Paper on Hummingbird+: low-cost FPGAs for LLM inference] Qwen3-30B-A3B Q4 at 18 t/s token-gen, 24GB, expected $150 mass production cost

A Hummingbird+ paper claims low-cost FPGAs run Qwen3-30B-A3B Q4 at 18 t/s generation. The title lists 24GB memory and an expected $150 mass-production cost; the post does not disclose FPGA model, power, or test conditions.

Why it matters: HKR-H/K/R all pass: the hook is a $150 FPGA running a 30B Q4 model, with speed, memory, and cost stated. Power, FPGA SKU, and test conditions are missing, so this lands at 79, not P1.

Read the original ↗Export Markdown