Skip to content
r/LocalLLaMA

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060. Here’s the mechanistic autopsy of what broke and what worked.

The author trained Mamba-Titan-1.4B-Reasoning on a 12GB RTX 3060: a frozen 1.4B Mamba-1 backbone with 8 trainable MoE arms, 2.54B total parameters, Top-2 routing at layers 24/25, and about 50% math accuracy.

Why it matters: HKR-H/K/R all pass via a numbered first-person experiment, but it is a single Reddit post with no independent replication and a fairly technical setup, so it stays in the low featured band.

Read the original ↗Export Markdown