Skip to content
Hacker News front page

Cactus taught Gemma 4 to know when it's wrong, routing only 15–35% of queries to a cloud model

Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong

Cactus post-trained Gemma 4 E2B to output a confidence score (0–1) with every response. High-confidence queries stay on-device; low-confidence ones are handed off to Gemini 3.1 Flash-Lite. Routing only 15–35% of queries matches the cloud model on most benchmarks. Instead of text self-rating or token entropy, they added a 68k-param probe layer that reads one intermediate hidden state and predicts p(wrong). Across 12 hold-out benchmarks the probe averages 0.814 AUROC vs 0.549 for token entropy. It was trained on zero audio data yet scores 0.79–0.88 on four audio benchmarks, suggesting it reads a modality-independent correctness signal. Weights are on HuggingFace, code is MIT-licensed, and it runs on Transformers, MLX, and Llama.cpp. The post doesn't disclose latency or on-device inference cost.

Why it matters: A concrete engineering solution with numbers, comparisons, and a deployment path—not concept hype. Not scoring higher because only the README is available; lacks large-scale eval data and failure cases. Real-world performance needs community validation.

Read the original ↗Export Markdown