The Artificial PostAccount
All papers
Chips · ORIGINAL RESEARCH

FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

FlashAttention-2 further improves GPU attention efficiency through work partitioning and parallelism.

Opening the paper…

KEEP FOLLOWING THE IDEA

Meet the researchers.

Tri Dao