闪电注意力

原名:flash-attention

加速长序列 Transformer 的训练和推理。

分类
开发提效
版本
v1.0.1
作者
弈韬(@ra1nzzz)
下载
2
收藏
0
发布
2026-08-25
更新
2026-09-29
TRACE 评分
3.2 / 5

内容概览

Flash Attention provides 2-4x speedup and 10-20x memory reduction for transformer attention through IO-aware tiling and recomputation. PyTorch native (easiest, PyTorch 2.2+) : flash-attn library (more features) : Copy this checklist: Step 1: Check PyTorch version If <2.2, upgrade: Step 2: Enable Flash Attention backend Replace standard attention: Force Flash Attention backend (torch.backends.cuda.sdp kernel is deprecated; use torch.nn.attention.sdpa kernel with SDPBackend): Step 3: Verify speedup with profiling Expected: 2-4x speedup for sequences 512 tokens. Step 4: Test accuracy matches baseline For multi-query attention, sliding window, or H100 FP8. Copy this checklist: Step 1: Install flash-attn library Step 2: Modify attention code Step 3: Enable advanced features Multi-query attention (shared K/V across heads): Sliding window attention (local attention): Step 4: Benchmark performan…

查看 SKILL 详情