Expert parallelism, two all-to-alls a layer, and why decoding a sparse model is harder than decoding a dense one