Attention Is All You Need
The transformer replaces recurrence with attention. It became a foundation for language models and many other systems that process sequences.
Paper & contextUsing attention to process sequences in parallel.
Vaswani studies architectures for learning from sequences and other structured data. The transformer is the starting point for this selection, followed by research on attention, efficient training, visual models, and large-scale pretraining data.
10 papers
The transformer replaces recurrence with attention. It became a foundation for language models and many other systems that process sequences.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources