flashattention FlashAttention Algorithm Solves Transformer Memory Bottlenecks A new IO-aware exact attention algorithm called FlashAttention speeds up Transformer training by optimizing memory reads and writes between GPU memory levels.