Researchers find that current large language models are severely undertrained, proving that model size and training data must scale equally for optimal performance. The resulting Chinchilla model outperforms massive rivals like GPT-3 using a fraction of the computing power.