trl-fine-tuning
TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.
- 分类
- 开发提效
- 版本
- v1.0.1
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 1
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-08-25
- TRACE 评分
- 3.2 / 5
内容概览
TRL provides post-training methods for aligning language models with human preferences. Installation : Supervised Fine-Tuning (instruction tuning): DPO (align with preferences): Complete pipeline from base model to human-aligned model. Note (TRL 1.x): PPO has been removed from TRL — PPOTrainer, PPOConfig, and python -m trl.scripts.ppo no longer exist. Use an online-RL trainer TRL still ships: RLOO (RLOOTrainer / trl rloo) is the closest drop-in for a reward-model-driven RLHF pipeline, and GRPO (GRPOTrainer / trl grpo, see Workflow 3) is the memory-efficient alternative. The step below uses RLOO. Copy this checklist: Step 1: Supervised fine-tuning Train base model on instruction-following data: Step 2: Train reward model Train model to predict human preferences: Step 3: RLOO reinforcement learning Optimize policy using the reward model. PPO was removed in TRL 1.x; use the RLOO CLI (trl rl…