torchtitan

Pretrain LLMs at scale with PyTorch 4D parallelism.

分类
开发提效
版本
v1.0.1
作者
弈韬(@ra1nzzz)
下载
1
收藏
0
发布
2026-08-25
更新
2026-08-25
TRACE 评分
3 / 5

内容概览

TorchTitan is PyTorch's official platform for large-scale LLM pretraining with composable 4D parallelism (FSDP2, TP, PP, CP), achieving 65%+ speedups over baselines on H100 GPUs. Installation : Download tokenizer : Start training on 8 GPUs : Copy this checklist: Step 1: Download tokenizer Step 2: Configure training In torchtitan's current layout, run configs are defined in a Python config registry (torchtitan/models/llama3/config registry.py) and selected by name via CONFIG=<name (or --config <name ). To customize, register your own config in the registry, or override individual fields on the command line (e.g. --optimizer.lr 3e-4 --training.steps 1000). The equivalent settings for an 8B run look like this (shown as fields; set them in the registry entry or as --section.key value overrides): Step 3: Launch training Step 4: Monitor and checkpoint TensorBoard logs are saved to ./outputs/tb…

查看 SKILL 详情