Inference 3 LLM Architecture Refresh [4]: Quantization, and Why Perplexity Won't Tell You It Broke Sep 5, 2026 LLM Architecture Refresh [3]: Flash Attention Is Exact, and Here's the Proof Sep 5, 2026 LLM Architecture Refresh [2]: The KV Cache, and Why Decode Is Memory-Bound Aug 16, 2026