# MoE 推理小技巧：不训练不微调，每 token 多激活几个专家，推理 token 数直接降 8.5%

> 原标题：Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!

- 来源：r/LocalLLaMA
- 发布时间：2026-09-03T21:58:56.000Z
- AX AI 日报：https://ai-daily.ax0x.ai/items/54626
- 原文：https://www.reddit.com/r/LocalLLaMA/comments/1w6lk6z/increasing_active_parameters_per_token_in_moe/

## 摘要

有人在 Qwen 35B A4B+ 上试了个 MoE 推理加速方法：每生成一个 token 时多激活几个专家参数，结果推理 token 数少了 8.5%，而且完全不用重新训练或微调。正文没披露具体怎么改的（比如改路由阈值还是硬选专家），也没说输出质量有没有掉。如果是真的，等于白捡一个加速，对部署 MoE 模型的人来说挺实用。
