SCALED DOT-PRODUCT ATTENTION
Attention(Q,K,V) = softmax(QKᵀ / √d_k)·V — each position gathers a weighted mix of value vectors based on query-key similarity.
CAUSAL MASK
Decoder self-attention blocks future positions — position i can only attend to positions ≤ i, so generation stays autoregressive.
RESIDUAL NORMALIZATION
Pre-LN (default): Normalizes input before the sub-layer, then adds output to residual. Trains from scratch without warmup. Post-LN: Adds first, then normalizes. The original Transformer architecture, but requires warmup or careful initialization to avoid gradient issues. Use the selector above to switch.
⚠ HAND-ROLLED AUTOGRAD
Every matmul, softmax, layernorm, and dropout gradient is derived and coded by hand — no autodiff library.