FlashAttention: Exact Attention as an IO-Aware Streaming Computation FlashAttention: IO 인식 스트리밍 계산으로서의 정확한 어텐션
FlashAttention is not an approximation to attention. Its core idea is to avoid materializing the N by N attention matrix in HBM by computing tiled attention in SRAM and maintaining online softmax statistics. FlashAttention은 어텐션의 근사가 아니다. 핵심은 SRAM에서 타일 단위 어텐션을 계산하고 online softmax 통계를 유지함으로써 HBM에 N by N 어텐션 행렬을 만들지 않는 것이다.
Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 4 - LLM Training Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 4 - LLM Training