← Flash PapersFlashAttention-2 Explained: Faster Attention with Better Parallelism and Work Partitioning

Share this Flash Paper

FlashAttention-2 Explained: Faster Attention with Better Parallelism and Work Partitioning

https://flashpapers.ai/p/s37ryspt38
Make your own
Details ▸
Forward Pass Peak Throughput:73%· A100 GPUBackward Pass Peak Throughput:63%· A100 GPUEnd-to-End Training Speed:225 TFLOPs/s· GPT-3 2.7B on 8xA100