Block AttnRes Architecture

24 layers divided into 3 blocks -- attention applied between blocks only

Within-block residual
Cross-block attention
Block boundary
Full AttnRes
O(L d) memory
L = 24 layers stored
Block AttnRes
O(N d) memory
N = 3 blocks, N << L