Long sequences and sparse models

Context parallelism and expert parallelism: the last two axes, and the two collectives that pay for them