A starting point.
Switch Transformers studies sparse mixture-of-experts models, activating selected expert networks rather than every parameter.
What to keep in mind
Read the original paper for the experimental setting, baselines, and limitations; a result in one setting is not a guarantee of performance elsewhere.
This work is included in a researcher’s reading path. A detailed editorial explanation is still being prepared. The complete manuscript is available in Full paper.
Source: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. The original manuscript contains the methods, experiments, figures, and references. An arXiv posting date may follow an earlier conference publication. Read the linked record for version history.