LLM Architecture Refresh 3 LLM Architecture Refresh [3]: Flash Attention Is Exact, and Here's the Proof Aug 2, 2026 LLM Architecture Refresh [2]: The KV Cache, and Why Decode Is Memory-Bound Aug 2, 2026 LLM Architecture Refresh [1]: Attention, the sqrt(d_k) Scale, and RoPE Aug 1, 2026