Compute-optimal is not cluster-optimal: MoE architecture choice depends on real cluster throughput
Compute-Optimal Is Not Cluster-Optimal
Sheng Zha's new paper MOSAIC folds system performance into scaling laws. Traditional workflow picks architecture by FLOPs budget first, then hands it to systems tuning—but clusters bill GPU-hours, not FLOPs. After fitting a joint scaling law on ~150 MoE pretraining runs, the paper finds: under FLOPs budget alone, higher sparsity always wins, pushing the optimum to the search boundary. On a real 512-GPU cluster, the sparsest design is the slowest—1.70× wall-clock vs. the densest. MOSAIC replaces model-FLOPs budget with deliverable FLOPs, co-optimizing parallel layout and memory constraints. A worked example: on four p6-B200 nodes for five days, configurations past ~0.96 sparsity cannot deliver the FLOPs their own recipe requires—the boundary optimum is infeasible.
Why it matters: Sheng Zha bakes system overhead into scaling laws and shows with ~150 MoE runs that 'theoretically optimal sparsity' can be infeasible on real clusters. A pragmatic correction to pretraining cost accounting, not pure theory flexing. Score stays below 85 because only the blog s...