A starting point.
The Vision Transformer applies a transformer to sequences of image patches and studies large-scale image recognition.
What to keep in mind
Read the original paper for the experimental setting, baselines, and limitations; a result in one setting is not a guarantee of performance elsewhere.
This work is included in a researcher’s reading path. A detailed editorial explanation is still being prepared. The complete manuscript is available in Full paper.
Source: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. The original manuscript contains the methods, experiments, figures, and references. An arXiv posting date may follow an earlier conference publication. Read the linked record for version history.