All the weights stay. Only some run.
One bank of weights. A different selection for each token.
One square = one expert FFN = 44.04 million weights.
Total · stored
11.32B257 FFNs + router ·
Active · one token
398.2M9 FFNs + router ·
3.5% of this sublayer’s weights are active. The rest stay stored.FFN = feed-forward network. Both counts include the router, but exclude attention.Same scale for both bars: 0–11.32B weights. M = million; B = billion. Displayed values are rounded.
Try another token: the selected FFNs change, but the parameter counts do not.
58 banks, plus the rest.
Add attention, three dense FFNs, and the input/output matrices. These published totals don’t change with the sliders.
Active weights are a subset, not a second model.
The equations, model breakdown & assumptions
- N
- Available routed expertsThe size of the stored bank.
- k
- Selected routed expertsHow many run for one token.
- E
- Weights in one expert FFN3 × 7,168 × 2,048 = 44,040,192.
Three SwiGLU weight matrices. - R
- Weights in the router7,168 × N.
Scores all routed experts for each token.
The in both equations is the shared expert: stored once, used for every token.
| Main-model component | Stored | Active / token |
|---|---|---|
| Sum of these matrices |
These are approximate, shape-based counts of the main prediction model’s matrices. They exclude normalization weights, router correction biases, quantization metadata and the auxiliary multi-token-prediction module. The input embedding contributes one row per token to the active count; the output head contributes its full matrix. This gives about 671.0B stored and 36.6B active, consistent with the report’s rounded 671B / 37B headline—not an exact checkpoint census.
Counts are calculated before rounding. The bars use a fixed linear scale, whose full length is one 256-expert V3 layer. Token selections are illustrative, not the routes V3 would choose for these words. The sliders keep one shared expert and the expert width fixed. Increasing available experts also increases the small router; it does not increase the number of expert FFNs evaluated when k stays fixed.
Stored weights drive weight-memory requirements. Active weights help estimate matrix-multiplication work, not wall-clock speed or the full training bill. Attention, batching, activations, optimizer states and communication still matter; inactive weights do not disappear from storage.
Sources: published configuration · DeepSeek-V3 technical report.