多卡加速
原名:accelerate
以最小改动实现 PyTorch 多 GPU 训练。
- 分类
- 开发提效
- 版本
- v1.0.1
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 2
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-09-29
- TRACE 评分
- 3.2 / 5
内容概览
Accelerate simplifies distributed training to 4 lines of code. Installation : Convert PyTorch script (4 lines): Run (single command): Original script : With Accelerate (4 lines added): Configure (interactive): Questions : - Which machine? (single/multi GPU/TPU/CPU) - How many machines? (1) - Mixed precision? (no/fp16/bf16/fp8) - DeepSpeed? (no/yes) Launch (works on any setup): Enable FP16/BF16 : Enable DeepSpeed ZeRO-2 (pass a DeepSpeedPlugin, not a raw dict): Or point at a full DeepSpeed JSON config via the plugin : ds config.json (a raw DeepSpeed config — passed via the plugin, NOT via --config file): Or via interactive config : Launch (--config file expects an accelerate YAML, not a raw DeepSpeed JSON): Enable FSDP : Or via config : Accumulate gradients : Effective batch size : batch size num gpus gradient accumulation steps Use Accelerate when : - Want simplest distributed training -…