Sina's VibeThinker-3B shows reasoning compresses into a 3B model, but factual knowledge doesn't
新浪开源VibeThinker-3B:推理可压缩,事实知识不能
Weibo's VibeThinker-3B, a 3B-parameter model, matches DeepSeek V3.2 and Kimi K2.5 on math and coding benchmarks despite being 200–333× smaller. Built on Alibaba's Qwen2.5-Coder-3B, it relies on multi-stage post-training. On knowledge-heavy GPQA-Diamond, it falls far behind large models. The team's takeaway: structured reasoning compresses well into small models; broad factual knowledge still needs scale.
Why it matters: Sina's VibeThinker-3B matches DeepSeek V3.2 and Kimi K2.5 on math and coding benchmarks, with disclosed training details and a useful finding that reasoning compresses well but factual knowledge doesn't. Not scored higher because only one source so far, and the model hasn't be...