A starting point.
CLIP learns visual concepts from paired images and text, connecting language descriptions to image representations.
What to keep in mind
Read the original paper for the experimental setting, baselines, and limitations; a result in one setting is not a guarantee of performance elsewhere.
This work is included in a researcher’s reading path. A detailed editorial explanation is still being prepared. The complete manuscript is available in Full paper.
Source: Learning Transferable Visual Models From Natural Language Supervision. The original manuscript contains the methods, experiments, figures, and references. An arXiv posting date may follow an earlier conference publication. Read the linked record for version history.