← Flash Papers
Training Compute-Optimal Large Language Models
Share
Share this Flash Paper
Training Compute-Optimal Large Language Models
https://flashpapers.ai/p/nckzxt2kkn
Copy link
Post to X
Post to Bluesky
Share on LinkedIn
Submit to Hacker News
Submit to Reddit
Email a link
Source PDF ↗
Open standalone ↗
Make your own
Make your own paper
Jordan Hoffmann, Sebastian Borgeaud +20 more
cs.CL
2022
arXiv:2203.15556
Code
· 3.2k★
· updated 2y ago
Details ▸
Details ▴
MMLU 5-shot accuracy:
67.6%
· MMLU
LAMBADA zero-shot accuracy:
77.4%
· LAMBADA
Wikitext103 Perplexity:
7.16
· Wikitext103
Winograde accuracy:
74.9%
· Winograde
MMLU 5-shot accuracy:
67.6%
· MMLU
LAMBADA zero-shot accuracy:
77.4%
· LAMBADA
Wikitext103 Perplexity:
7.16
· Wikitext103
Winograde accuracy:
74.9%
· Winograde
Concepts:
Transformer