← Flash PapersSwitch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Share this Flash Paper

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

https://flashpapers.ai/p/gx625kuhf4
Make your own
William Fedus, Barret Zoph, Noam Shazeer
Details ▸
pre-training speedup:7x· Colossal Clean Crawled Corpus (C4)parameter scale:1.6 trillionpre-training speedup:4x· C4quality gain retention on distillation:30%· SuperGLUE