闪电注意力
原名:flash-attention
加速长序列 Transformer 的训练和推理。
- 分类
- 开发提效
- 版本
- v1.0.1
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 2
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-09-29
- TRACE 评分
- 3.2 / 5
内容概览
Flash Attention provides 2-4x speedup and 10-20x memory reduction for transformer attention through IO-aware tiling and recomputation. PyTorch native (easiest, PyTorch 2.2+) : flash-attn library (more features) : Copy this checklist: Step 1: Check PyTorch version If <2.2, upgrade: Step 2: Enable Flash Attention backend Replace standard attention: Force Flash Attention backend (torch.backends.cuda.sdp kernel is deprecated; use torch.nn.attention.sdpa kernel with SDPBackend): Step 3: Verify speedup with profiling Expected: 2-4x speedup for sequences 512 tokens. Step 4: Test accuracy matches baseline For multi-query attention, sliding window, or H100 FP8. Copy this checklist: Step 1: Install flash-attn library Step 2: Modify attention code Step 3: Enable advanced features Multi-query attention (shared K/V across heads): Sliding window attention (local attention): Step 4: Benchmark performan…