Skip to content
Hacker News front page

Fermion ships Neutrino-1 8B: one 3.88 GB file that runs on H100, L4, and MacBook

Neutrino-1 8B

Fermion Research released Neutrino-1 8B, an 8.19B-parameter dense decoder-only model. Its 252 linear layers use a proprietary ternary-family weight format, shrinking the whole model to 3.88 GB—one-eighth of fp16. Weights stay bit-packed on disk and are decoded inside matrix kernels, so no fp16 or fp32 weight material lives in the decode path. The small working set changes serving economics: single-stream decode is bound by bytes moved per token, so 3.88 GB decodes faster than a 16 GB fp16 artifact on the same memory system, and the model fits on an 8 GB GPU or a 16 GB laptop. One container runs on datacenter GPUs, MacBooks, and desktop CPUs without conversion. MMLU 5-shot hits 72.1. On H100, plain decode reaches 396 tok/s and speculative decode hits 763 tok/s; an M5 MacBook gets 33.7 tok/s via MLX and 24.9 tok/s CPU-only. The model is based on Qwen3-8B under Apache-2.0. The post does not disclose training data or method details.

Why it matters: Neutrino-1 8B compresses an 8.19B-param model to 2.56 GB — 1/8 of fp16 — with 72.1 MMLU and 763 tok/s speculative decode throughput. The technical detail is concrete, not marketing fluff. Held at 78 because it's a debut from a new lab with no third-party reproduction or cross-...

Read the original ↗Export Markdown