What an MoE layer costs to run

Expert parallelism, two all-to-alls a layer, and why decoding a sparse model is harder than decoding a dense one