Skip to content
r/LocalLLaMA

MoE inference trick: activate more params per token, cut reasoning tokens by 8.5% without retraining

Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!

A Reddit user found that increasing active parameters per token in MoE models like Qwen 35B A4B+ reduces reasoning tokens by 8.5% without any training or fine-tuning. The post doesn't explain how to implement the change or whether output quality holds, but the idea is a practical free speed-up for MoE deployments.

Read the original ↗Export Markdown