← Flash Papers
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Share
Share this Flash Paper
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
https://flashpapers.ai/p/gx625kuhf4
Copy link
Post to X
Post to Bluesky
Share on LinkedIn
Submit to Hacker News
Submit to Reddit
Email a link
Source PDF ↗
Open standalone ↗
Make your own
Make your own paper
William Fedus, Barret Zoph, Noam Shazeer
cs.LG
2021
arXiv:2101.03961
Code
· 3.0k★
· updated 2mo ago
Details ▸
Details ▴
pre-training speedup:
7x
· Colossal Clean Crawled Corpus (C4)
parameter scale:
1.6 trillion
pre-training speedup:
4x
· C4
quality gain retention on distillation:
30%
· SuperGLUE
pre-training speedup:
7x
· Colossal Clean Crawled Corpus (C4)
parameter scale:
1.6 trillion
pre-training speedup:
4x
· C4
quality gain retention on distillation:
30%
· SuperGLUE
Concepts:
Mixture of Experts