Skip to content
Hacker News front page

Cactus Needle 3: 8-29 MB automation models beat DeepSeek V4 Flash on tool calls

Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash

Cactus open-sourced Needle 3, an automation model for phones, wearables, robots, and other tiny devices. The whole model is 8-29 MB, built on Simple Attention Networks where every layer is a usable sub-network. A fine-tuned 4-layer sub-network beats DeepSeek V4 Flash on mobile tool-calling benchmarks; on extraction it matches models 2-3x its size. It runs fully offline at 400-4,000 tokens/s decode on a Raspberry Pi 5. The Python package supports tool calls, structured extraction, and text embeddings. The post does not disclose training data composition or fine-tuning cost.

Why it matters: A tiny model claims to match DeepSeek V4 Flash on mobile tool calls at 8-29MB, with a novel intelligence ladder architecture. Score capped below 85 because only the project page is available—no third-party benchmarks or production deployment stories to cross-validate the perfo...

Read the original ↗Export Markdown