Skip to content
r/LocalLLaMA

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Cactus Compute open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, using an attention-and-gating architecture with no MLPs.

Why it matters: HKR-H/K/R all pass: a 26M tool-calling model has a strong hook and concrete speed/design claims. Single Reddit source and a less-known team keep it in the lower 78–84 band.

Read the original ↗Export Markdown