FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
FlashAttention reorganizes attention computation to reduce costly reads and writes between GPU memory levels.
Paper & contextMaking learning systems more efficient and usable.
Ré studies the systems, data, and algorithms that make machine learning practical. These papers connect efficient model execution with language-model evaluation and data-centric development, treating system design as part of the learning problem.
10 papers
FlashAttention reorganizes attention computation to reduce costly reads and writes between GPU memory levels.
Paper & contextSelected research in language models, ai evaluation. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in computer vision, language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources