Two axes of a (T, d) matrix

Attention moves information between rows with weights it computes on the fly. The FFN rewrites each row on its own using weights it learned once and reuses for every token.