Training Compute-Optimal Large Language Models
The Chinchilla study revisits how to divide a training budget between model size and data. It argues that many contemporary large models were undertrained.
Paper & contextStudying compute-efficient training and sparse language models.
Mensch’s coauthored research spans compute-efficient language-model training and multimodal models. Chinchilla and Mixtral are starting points for this selection, followed by work on open models, coding, speech, and reasoning.
10 papers
The Chinchilla study revisits how to divide a training budget between model size and data. It argues that many contemporary large models were undertrained.
Paper & contextMixtral of Experts explores a sparse mixture-of-experts architecture, using a subset of experts for each token.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models, computer vision. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in ai for science, reinforcement learning. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources