The Artificial PostAccount
All papers
Chips · ORIGINAL RESEARCH

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

FlashAttention reorganizes attention computation to reduce costly reads and writes between GPU memory levels.

Opening the paper…

KEEP FOLLOWING THE IDEA

Meet the researchers.

Stefano ErmonTri DaoChristopher Ré