Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax
Luce Spark runs Qwen3.6 35B-A3B at 13.3 GiB peak VRAM on an RTX 3090 by keeping hot experts on GPU, swapping cold experts through a bounded async cache, and using one fused graph for decode at about 100 tok/s.
Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE on a 16 GB GPU, with 13.3 GiB peak use and ~100 tok/s. Reddit-source and no third-party replication keep it at 78.