FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
FlashAttention reorganizes attention computation to reduce costly reads and writes between GPU memory levels.
Paper & contextReducing memory traffic and computation in sequence models.
Dao studies the relationship between learning algorithms and the hardware that runs them. FlashAttention and Mamba anchor this selection, connecting memory-efficient attention with sequence models designed to use computation and memory differently.
10 papers
FlashAttention reorganizes attention computation to reduce costly reads and writes between GPU memory levels.
Paper & contextMamba investigates selective state-space models as an alternative approach to sequence modeling with efficient scaling.
Paper & contextFlashAttention-2 further improves GPU attention efficiency through work partitioning and parallelism.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources