torchtitan
Pretrain LLMs at scale with PyTorch 4D parallelism.
- 分类
- 开发提效
- 版本
- v1.0.1
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 1
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-08-25
- TRACE 评分
- 3 / 5
内容概览
TorchTitan is PyTorch's official platform for large-scale LLM pretraining with composable 4D parallelism (FSDP2, TP, PP, CP), achieving 65%+ speedups over baselines on H100 GPUs. Installation : Download tokenizer : Start training on 8 GPUs : Copy this checklist: Step 1: Download tokenizer Step 2: Configure training In torchtitan's current layout, run configs are defined in a Python config registry (torchtitan/models/llama3/config registry.py) and selected by name via CONFIG=<name (or --config <name ). To customize, register your own config in the registry, or override individual fields on the command line (e.g. --optimizer.lr 3e-4 --training.steps 1000). The equivalent settings for an 8B run look like this (shown as fields; set them in the registry entry or as --section.key value overrides): Step 3: Launch training Step 4: Monitor and checkpoint TensorBoard logs are saved to ./outputs/tb…