Skip to content
Hacker News front page

DeepSeek-V4 Latent Reasoning ships as a self-contained model, not an adapter

Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space

Nicholai Mitchko turned the CoLaR latent reasoning head into a single deployable model. It uses a DeepSeek-V4-Flash-0731 backbone quantized to NVFP4 (~79 GiB/GPU at TP=2) and a 35.7M-param reasoning head. BBH zero-shot aggregate is 0.94, with perfect scores on multi-step state tracking but 0.26 on Dyck languages. A forked vllm runtime serves it, with per-request reasoning depth control via HTTP headers.

Why it matters: Turning CoLaR latent reasoning from an external adapter into a full model has engineering value, and BBH zero-shot 0.94 is solid. But it's a personal research blog with no cross-source verification and a narrow audience, so it lands right at the featured threshold.

Read the original ↗Export Markdown