attention 3 LLM Architecture Refresh [3]: Flash Attention Is Exact, and Here's the Proof Sep 5, 2026 LLM Architecture Refresh [1]: Inside a Transformer Block — Attention, Heads, and the FFN Aug 15, 2026 Study Notes: Stanford CS336 Language Modeling from Scratch [5] Sep 13, 2025