0:00 / 0:22

Swipe to inspect the full diagram

This illustrative two-feature model uses four 2→3→2 ReLU feed-forward networks and a learned top-2 router. The FFN boundary remains fixed inside a schematic transformer block; normalization is omitted from the outer block for space. The initial dense-to-MoE change illustrates architectural replacement, not a weight-preserving split. Only selected experts evaluate their up projection, activation and down projection. Their full output vectors are weighted and added. The model is not trained or a DeepSeek implementation. The router and expert bank replace the dense FFN; attention is unchanged. Four experts are stored but two evaluate per token. Router work, storage and communication still have costs. Colors denote signed scalar values: green positive, brown negative. Choose among four illustrative token vectors. Drag the timeline or use the operation controls to inspect any step. Playback stops at the result.