← Flash PapersFlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Share this Flash Paper

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

https://flashpapers.ai/p/ahd5qyqmjy
Make your own
Tri Dao, Daniel Y. Fu +3 more
Details ▸
BERT-large End-to-End Training Speedup:15%· WikipediaGPT-2 Training Speedup:3.0x· OpenWebTextPath-X Accuracy (seq length 16k):61.4%· Long Range ArenaPath-256 Accuracy (seq length 64k):63.1%· Long Range ArenaMemory Footprint Reduction:up to 20x· LRA (seq length 16k)