Learning Transferable Visual Models From Natural Language Supervision
CLIP learns visual concepts from paired images and text, connecting language descriptions to image representations.
Paper & contextLearning transferable representations through large-scale pretraining.
Radford’s research explores what models can learn from large and varied datasets. These papers connect generative visual models, language pretraining, image-text representations, and speech recognition, with an emphasis on representations that transfer across tasks.
10 papers
CLIP learns visual concepts from paired images and text, connecting language descriptions to image representations.
Paper & contextGPT-3 investigates how a large language model can perform tasks from instructions and examples in its prompt, without task-specific weight updates.
Paper & contextSelected research in multimodal ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextHow model size, training data and compute relate to language-model performance. A foundation for understanding scaling as an empirical relationship, with limits.
Paper & contextSelected research in language models, multimodal ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextThe GPT-4 technical report documents capabilities, evaluations and limitations of a multimodal model, while withholding many architecture and training details.
Paper & contextProximal Policy Optimization proposes a policy-gradient objective that limits overly large updates while simplifying implementation.
Paper & contextSelected research in multimodal ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in multimodal ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models, computer vision. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources