多卡加速

原名:accelerate

以最小改动实现 PyTorch 多 GPU 训练。

分类
开发提效
版本
v1.0.1
作者
弈韬(@ra1nzzz)
下载
2
收藏
0
发布
2026-08-25
更新
2026-09-29
TRACE 评分
3.2 / 5

内容概览

Accelerate simplifies distributed training to 4 lines of code. Installation : Convert PyTorch script (4 lines): Run (single command): Original script : With Accelerate (4 lines added): Configure (interactive): Questions : - Which machine? (single/multi GPU/TPU/CPU) - How many machines? (1) - Mixed precision? (no/fp16/bf16/fp8) - DeepSpeed? (no/yes) Launch (works on any setup): Enable FP16/BF16 : Enable DeepSpeed ZeRO-2 (pass a DeepSpeedPlugin, not a raw dict): Or point at a full DeepSpeed JSON config via the plugin : ds config.json (a raw DeepSpeed config — passed via the plugin, NOT via --config file): Or via interactive config : Launch (--config file expects an accelerate YAML, not a raw DeepSpeed JSON): Enable FSDP : Or via config : Accumulate gradients : Effective batch size : batch size num gpus gradient accumulation steps Use Accelerate when : - Want simplest distributed training -…

查看 SKILL 详情