LLM Architecture Refresh 5
- LLM Architecture Refresh [5]: Mixture-of-Experts, and Why Sparsity Doesn't Survive a Batch
- LLM Architecture Refresh [4]: Quantization, and Why Perplexity Won't Tell You It Broke
- LLM Architecture Refresh [3]: Flash Attention Is Exact, and Here's the Proof
- LLM Architecture Refresh [2]: The KV Cache, and Why Decode Is Memory-Bound
- LLM Architecture Refresh [1]: Inside a Transformer Block — Attention, Heads, and the FFN